Multiplex sample analysis method and apparatus, storage medium, and device
By constructing a pre-defined classifier to identify and filter misallocated reads in multiplex sequencing, the problem of false positives and false negatives caused by label skipping is solved, thus improving the accuracy of sample detection.
Patent Information
- Application Number
- PCT/CN2025/084070
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-19
- Filing Date
- 2025-03-21
- Publication Date
- 2025-10-23
AI Technical Summary
Label skipping in multiplex sequencing increases false positives and false negatives, which is particularly problematic in applications such as cancer genomics and liquid biopsy that require precise detection of rare variants.
By constructing a pre-defined classifier and utilizing the feature information in the sequencing data, misallocated reads that may cause tag skipping can be identified and filtered out. By employing feature quantization and signal analysis, the classification results of reads can be accurately predicted, thereby reducing the impact of misallocation.
It effectively reduces misassignment caused by label skipping, improves the accuracy and reliability of sample detection, and reduces false positive and false negative results.
Smart Images

Figure CN2025084070_23102025_PF_FP_ABST
Abstract
Description
A multiple sample analysis method, device, storage medium and apparatus TECHNICAL FIELD
[0001] The present application relates to the technical field of bioinformatics, and more particularly to a multiple sample analysis method, device, storage medium and apparatus. BACKGROUND
[0002] Next-generation sequencing, also known as high-throughput sequencing or massively parallel sequencing, can determine nucleic acid sequences of multiple samples in one sequencing run. Such determination can be achieved by multiple sample analysis, also known as multiplex library or multiplex sequencing. In multiplex sequencing, a specific sequence corresponding to a sample uniquely is added to each nucleic acid fragment during library construction, so that the libraries of multiple samples can be mixed in one reaction system for sequencing to obtain sequencing data, and then the sequencing data can be allocated to the corresponding sample according to the specific sequence, so as to obtain the sequencing data of each sample. The specific sequence is usually referred to as a tag or index (barcode or tag).
[0003] Tag misassignment between multiplex libraries, also known as index hopping or index misassignment or sample cross-talk, is a problem in multiplex sequencing. Understandably, any detection involving the use of high-throughput sequencing to seek trace "positive" data in a high-background noise interference mixture is very susceptible to index hopping, including cancer genomics and other applications requiring accurate detection of rare variants, such as liquid biopsy, pathogen detection, etc. For example, in pathogen detection applications, even if a single-digit number of reads of a pathogen is detected, it is generally considered as a positive detection of the pathogen. In this case, if the reads of one pathogen are mistakenly divided into another pathogen, i.e., the tag misassignment causes the reads from pathogen A to be mistakenly allocated to another pathogen B, this will lead to an increase in false positives or false negatives.
[0004] It can be seen that reducing index hopping and reducing misassignment are currently worth paying attention to problems in sequencing applications. SUMMARY
[0005] Therefore, embodiments of the present application provide a multiple sample analysis method, a multiple sample analysis device, a method of constructing a classifier capable of predicting read categories, a terminal or computing device, a computer-readable storage medium, etc.
[0006] The methods of building classifiers, multiplex sample analysis methods, etc. of the embodiments of the present application are made by the inventors after observing and comparing sequencing data of a large number of cases of label jumping and cases of no label jumping, and the following conclusions and findings are made:
[0007] Common multiplex sample analysis based on surface fluorescence imaging sequencing technology has established specific correspondence between indexes and samples before mixing the libraries of each sample, and each library contains indexes and inserts (nucleic acid fragments of the sample to be detected, often also referred to as inserts, to-be-detected fragments, to-be-sequenced sequences, inserts, or insert ions in library construction), so after obtaining the sequencing data of the mixed samples, the associated insert fragment reads are usually read out according to the indexes to be assigned to the corresponding sample to obtain the sequencing data of the specified sample for detection and analysis of the specified sample. This process is also commonly referred to as read splitting or assignment or demultiplexing.
[0008] For a plurality of clusters (amplification products of the library) on the surface to be subjected to sequence determination, desirably or generally, a cluster includes a plurality of polynucleotide molecules of the same sequence, and sequencing the cluster can produce a group of reads of the same sequence (sometimes also referred to as the same kind of read below). By observing the conversion acquisition process such as the conversion of biochemical reaction signals into image signals and the final determination of bases based on image signals to determine reads, the inventors have found that, from the image information after acquisition and processing, from the signal changes and detection results of the specified clusters that have undergone biochemical reactions in the image, compared with the detection signals and identification results of the same to-be-detected part of a normal cluster, which mostly show “one cluster, one index, and one read”, the clusters corresponding to the reads that are incorrectly assigned due to label jumping can be classified into the following four categories: reading out one index and one read from a cluster, and the readout quality of the index and / or read is low, reading out multiple indexes and one read from a cluster, reading out one index and multiple reads from a cluster, and reading out multiple indexes and multiple reads from a cluster. Of course, the detection signals and identification results of the to-be-detected part of some clusters simultaneously exhibit multiple of the above situations. Moreover, the inventors have found after statistics that any of the latter three situations is more likely to cause the sequencing data to be incorrectly split to a non-true source sample.
[0009] Based on the summary findings, the inventors utilized the training data with normal reads (reads that were correctly assigned to their true source sample, sometimes also referred to herein as non-string reads) and abnormal reads (reads that were incorrectly assigned to a sample other than their true source sample, sometimes also referred to herein as mismatched reads or hopping- or jumping- or stringed reads) as classification labels, constructed, tested and screened a plurality of features of the reads, particularly those that were highly associated with the observed cases described above, and ultimately determined the features of the reads that were able to better distinguish or fit the distribution of the two types of labels to train the function relationship / model of the classification labels and features.
[0010] The model was verified to be able to accurately predict the classification results of the reads (whether they are mismatched reads or the probability of being mismatched reads or normal reads). For this purpose, the embodiments of the present application provide a multiplex sample analysis method, comprising:
[0011] obtaining sequencing data of a multiplex sample, the sequencing data comprising reads from a plurality of samples;
[0012] splitting the sequencing data to obtain reads assigned to a specified sample;
[0013] filtering out a corresponding part of reads belonging to a specified category of reads from the reads assigned to the specified sample according to a classification result of the reads assigned to the specified sample by a preset classifier, comprising obtaining values of features of the reads assigned to the specified sample, and querying the preset classifier according to the values of the features to determine the classification result of the reads assigned to the specified sample,
[0014] wherein the preset classifier is a quantification scheme associated with the features and the probability of the reads belonging to the specified category of reads, and the features are a plurality of quantifiable features generated during sequencing runs and related to signals for obtaining the reads assigned to the specified sample and / or assigning the reads to the specified sample;
[0015] and detecting the specified sample according to the remaining reads assigned to the specified sample.
[0016] A multiplex sample analysis device, comprising:
[0017] an obtaining unit configured to obtain sequencing data of a multiplex sample, the sequencing data comprising reads from a plurality of samples;
[0018] a splitting unit configured to split the sequencing data to obtain reads assigned to a specified sample;
[0019] a filtering unit configured to filter out a corresponding part of reads belonging to a specified category of reads from the reads assigned to a specified sample according to a classification result of the reads assigned to the specified sample by a preset classifier, including obtaining a value of a feature of the reads assigned to the specified sample, and querying the preset classifier according to the value of the feature to determine the classification result of the reads assigned to the specified sample,
[0020] wherein the preset classifier is a quantization scheme associated with the feature and a probability of the reads belonging to the specified category of reads, and the feature is a plurality of quantifiable features generated in a sequencing run and related to a signal for obtaining the reads assigned to the specified sample and / or assigning the reads to the specified sample;
[0021] and a detecting unit configured to detect the specified sample according to the remaining reads assigned to the specified sample.
[0022] A computer readable storage medium for storing a program executable by a computer, and executing the program includes completing the multiple sample analysis method according to any one of the above.
[0023] A computing device, comprising:
[0024] a memory configured to store a computer executable program;
[0025] one or more processors configured to execute the computer executable program and implement the multiple sample analysis method according to any one of the above.
[0026] A method for constructing a classifier capable of predicting whether reads assigned to a specified sample from mixed sample sequencing data are string reads or non-string reads, the method comprising:
[0027] obtaining a training data set, the training data set being sequencing data of multiple samples with known sample assignment results of reads and true value assignment results thereof, the training data set including a plurality of reads containing at least a part of a sequence to be sequenced of each sample and a specific label sequence carried by the sample, and the training data set containing at least a certain number of reads;
[0028] dividing the reads into two categories according to whether the assignment results and the true value assignment results are consistent, and quantifying each read in the two categories of reads into an alternative form based on a feature of the read;
[0029] counting the number of each category of reads contained in each alternative form to establish a corresponding relationship between the alternative form and the category of reads;
[0030] constructing a classification classifier based on the corresponding relationship.
[0031] A computer readable storage medium storing a program for execution by a computer, execution of the program comprising carrying out the method of building a classifier as described above.
[0032] A computer device comprising:
[0033] a memory for storing a computer executable program;
[0034] one or more processors for executing the computer executable program with implementing the method of building a classifier as described above.
[0035] A method of applying a classifier, comprising:
[0036] classifying reads assigned to a given sample to obtain a classification result;
[0037] such that, according to the classification result, a corresponding portion of reads belonging to a given class is filtered out, including obtaining values of features of the reads assigned to the given sample, and querying the pre-set classifier according to the values of the features to determine the classification result of the reads assigned to the given sample; and detecting the given sample according to the remaining reads assigned to the given sample;
[0038] wherein the reads assigned to the given sample are obtained by splitting sequencing data of multiple samples, and the sequencing data includes reads from multiple samples.
[0039] The multiple sample analysis method or device disclosed in the embodiments of the present application obtains the classification result of reads assigned to a given sample through a pre-set classifier and filters out a portion of reads belonging to a given class, such as those reads that are incorrectly assigned to the given sample due to a label jump with a relatively high probability, thereby reducing the impact of label jumps and incorrectly assigned reads on detection of the given sample. BRIEF DESCRIPTION OF DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort based on the provided drawings.
[0041] FIG. 1 is a flowchart of a multiple sample analysis method provided by an embodiment of the present application;
[0042] FIG. 2 is a structural diagram of a multiple sample analysis device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0043] The technical solutions in the embodiments of the present application will be described in detail below with reference to the drawings in the embodiments of the present application. The described embodiments are exemplary and are only used to explain the present application, and cannot be understood as a limitation of the present application.
[0044] In the description of the present application, "comprising" and "having" and any variants thereof cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, and can also include steps or units that are not listed but have similar features or functions or do not conflict with the listed steps or units or means.
[0045] Label assignment error between multiple libraries, also known as index hopping or index misassignment or sample cross-talk, is a problem in the process of multiplex sequencing. In order to solve this problem to some extent or reduce the influence of this problem on sample detection to some extent, the embodiments of the present application propose a multiplex sample analysis method. Please refer to FIG. 1, the method comprises: S101 obtaining sequencing data of a multiplex sample, the sequencing data comprising reads from multiple samples; S102 splitting the sequencing data to obtain reads assigned to a specified sample; S103 filtering out the corresponding part of reads belonging to the specified category of reads according to the classification result of the reads assigned to the specified sample by a preset classifier, comprising obtaining the values of the features of the reads assigned to the specified sample and determining the classification result of the reads assigned to the specified sample according to the values of the features by querying the preset classifier, wherein the preset classifier is a quantization scheme associated with the features and the probability of the reads belonging to the specified category of reads. The features referred to are a plurality of quantifiable features generated in the sequencing run and related to the signals for obtaining the reads assigned to the specified sample and / or assigning the reads to the specified sample, and then S104 detecting the specified sample according to the remaining reads assigned to the specified sample. This method can reduce the adverse effects of label hopping on sample detection.
[0046] Specifically, in step S101, sequencing data of a multiplex sample is obtained, and the sequencing data comprises reads from multiple samples. The sequencing data of the multiplex sample refers to sequencing data obtained by mixing libraries of multiple samples in one reaction system for sequencing.
[0047] In order to sequence a plurality of samples in one reaction, the samples from different sources need to be distinguished in advance, i.e. different tags (index or barcode) are added to samples from different sources, which needs to be completed in the process of library construction. For example, taking the library construction of the second-generation sequencing platform of Illumina Company as an example, the main process includes: breaking the sample, e.g. DNA genome, into DNA fragments according to the Illumina / Solex DNA sample preparation method, then repairing the sticky ends formed by breaking into blunt ends; then adding a base "A" to the 3' end so that the DNA fragments can be connected to the adaptor with "T" base at the 3' end and containing a tag sequence for marking the source of the sample; then using PCR technology to amplify the DNA fragments with adaptors at both ends and purifying the final PCR product to obtain a library containing "tag" and "insert" parts. The specific process of library construction can also refer to the literature Multiplexed Illumina sequencing libraries from picogram quantities of DNA, Bowman et al. BMC Genomics 2013, 14:466, http: / / www.biomedcentral.com / 1471-2164 / 14 / 466 and / or the literature Multiplex Illumina sequencing using DNA barcoding, Koon Ho Wong, Curr. Protoc. Mol. Biol. 101:7.11.1-7.11.11., 2013, https: / / currentprotocols.onlinelibrary.wiley.com / doi / 10.1002 / 04 71142727.mb0711s101, the contents of which are all incorporated herein. Among them, the multiplex library is a mixture of sequencing libraries of a plurality of samples, and each sequencing library of the sample has a tag corresponding specifically to the sample. In some examples, the tag is a known sequence of greater than or equal to 1 bp, and the tags corresponding specifically to any two samples have a difference of at least one base. Further, in some examples, the tag is a known sequence of greater than or equal to 4 bp and less than 20 bp, and the tags corresponding specifically to any two samples have a difference of at least three bases to improve the fault tolerance.
[0048] When sequencing, the prepared libraries of the multiplex samples are mixed in one reaction system and sequenced simultaneously. For example, in the edge-sequencing based on surface fluorescence microscopic imaging, the reaction system generally includes sequencing primers, polymerase, reaction substrates. Under the catalysis of DNA polymerase and under the conditions suitable for polymerization reaction, the (multiplex) library is combined with the sequencing primer (the library or the amplified library or the library-primer complex is also called the nucleic acid molecule to be tested or the template), and based on the principle of base pairing, the reaction substrates such as nucleotides or their analogs are controllably connected to the end of the sequencing primer or the template, and the base at the corresponding position of the library / template is paired and connected, so as to realize the extension of the sequencing primer. The nucleotides or their analogs connected to the end of the sequencing primer are also called extension bases. The extension bases can be provided with an optically detectable label such as a fluorescent group, and different types of bases incorporated into the template (extension bases) can be associated or distinguished by different fluorescent groups or different light-emitting wave bands of the fluorescent groups, so that the type of the extension base can be determined by collecting the light-emitting signals generated by the labels during the extension of the sequencing primer by the microscopic imaging system, so as to obtain the sequencing data. The sequencing of the multiplex samples can also refer to the document Multiplex Illumina sequencing using DNA barcoding, Koon Ho Wong, Curr. Protoc. Mol. Biol. 101:7.11.1-7.11.11., 2013 https: / / currentprotocols.onlinelibrary.wiley.com / doi / 10.1002 / 0471142727.mb0711s101 and the document https: / / www.illumina.com / documents / products / datasheets / datasheet_sequencing_multiplex.pdf and the document Nature. 2008 Nov 6; 456(7218): 53-59, https: / / www.ncbi.nlm.nih.gov / pmc / articles / PMC 2581791 / , the contents of which are all incorporated herein.
[0049] In some examples, at the time of sequencing, a template can be constructed from one or a few rounds of imaging, where the template as referred to herein corresponds to a template in a biochemical reaction, a set of locations of chemical features (templates) on the surface, e.g., molecular clusters (amplified library) in an image; based on the images formed from the light signals generated by the optically detectable labels during the sequencing primer extension process and the constructed template, the type of the extended base incorporated into the nucleic acid molecule under test in the round of reaction is identified to obtain sequencing data. The template construction and base identification process can be referred to PCT application WO2020037574A1, Chinese patent CN112823352B, and the literature Base-calling for next-generation sequencing platforms, Brief Bioinform, 2011 Sep; 12(5): 489-497, the contents of which are incorporated herein in their entirety.
[0050] In some examples, the multiplex library comprises the insert and the tag linked thereto, and the multiplex library is sequenced based on the surface fluorescence microscopic imaging detection for sequencing-by-synthesis, comprising:
[0051] linking the multiplex library to the surface and amplifying the multiplex library on the surface, or amplifying the multiplex library and loading the amplified multiplex library to the surface, to form the surface with a plurality of amplicons;
[0052] subjecting the surface to conditions suitable for polymerization reaction, and performing multiple rounds of sequencing on the amplicons, comprising acquiring images of each round of sequencing and detecting signals on the images corresponding to the locations of the amplicons, to identify the type of the base incorporated into the specified amplicon in the round of reaction, to determine the sequencing sequence of at least a portion of the insert and the sequencing sequence of the tag in the specified amplicon.
[0053] In certain embodiments, the amplification can be performed on the multiplex library in a solid phase or a liquid phase to obtain the amplified product of the multiplex library. Without limitation to the amplification method, the amplification can be performed by PCR (polymerase chain reaction) using Taq enzyme, or RPA (recombinase polymerase amplification) using Bst or Bsu or recombinase or multi-enzyme system, or isothermal amplification technology such as SDA (strand displacement amplification) using RCA (rolling circle amplification), and the like. For example, the amplification can be performed on the multiplex library on a solid phase surface by bridge PCR or template walking amplification to form amplicons on the surface; or the amplification can be performed on the multiplex library in a liquid phase by RCA to obtain the amplified product of the multiplex library, and then the amplified product is loaded onto the surface to form DNA nanoballs on the surface, and the like. (The multiplex library or the amplified multiplex library, without involving the need to distinguish the library before mixing, without affecting the understanding of those skilled in the art, is sometimes referred to as the library or template or nucleic acid molecule to be detected or amplified product or cluster or nanoball in this paper.)
[0054] In certain examples, when sequencing, the "tag" part and the "insert" part are not limited to which part is read first and which part is read later, nor are they limited to one-time reading (e.g., using one sequencing primer to read the tag and the corresponding insert) or multiple-time reading (e.g., using two or more sequencing primers to read one or more tags and the corresponding insert, respectively). For example, the library structure or one repeat unit of the library structure can be represented as 5' adapter-insert-tag-adapter, and at least one part of the tag and the insert can be read in sequence using a sequencing primer complementary to at least one part of the 3' adapter, and the tag part and the insert part in the read sequence can be distinguished according to the predetermined length of the tag and the tag-to-insert reading order. For another example, the library structure or one repeat unit of the library structure can be represented as 5' adapter-tag2-insert-tag1-adapter, and at least one part of the tag1 and the insert can be read in sequence using a sequencing primer 1 complementary to at least one part of the 3' adapter, and at least one part of the tag2 and the insert can be read using a sequencing primer 2 identical to the 5' adapter, and the tag (tag1 and tag2) part and the insert part can be determined based on the two read sequences. For another example, the library structure or one repeat unit of the library structure can be represented as 5' adapter-insert-X-tag-adapter, where X is a sequence-known spacer, and the tag sequence can be determined using a sequencing primer 1 complementary to at least one part of the 3' adapter, and a part of the insert can be determined using a sequencing primer 2 complementary to X. The specific process of sequencing can also be referred to in Chinese patent application CN113293205A, the contents of which are incorporated herein in their entirety.
[0055] The reads or readouts or readouts referred to herein, the sequencing sequences or sequencing sequences, a read is a sequencing readout of a polynucleotide sequence, sometimes refers to a readout containing index and insert (or part of the insert), sometimes refers to the readout of the insert (or part of the insert), sometimes refers to the readout of the index; in specific examples, those skilled in the art can clearly understand according to the relevant annotations and / or context. Generally, in multiplex sample detection analysis, if the sequencing data is split and assigned to each sample before, the reads generally contain tag readouts (tag readouts and corresponding insert readouts can be in a read, or in multiple reads that can determine the relative position relationship; the "corresponding insert" and "relative position relationship" referred to here means that the tag and the insert come from the same template), if the sequencing data is split and assigned to the specified sample for specified sample detection, then the reads generally refer to the insert readout without tag readout.
[0056] After obtaining the sequencing data of multiple samples, the sequencing data needs to be split according to step S102 to obtain the reads assigned to the specified sample. The sequencing data includes reads of multiple samples, so the sequencing data can be split to obtain reads assigned to the specified sample, which refers to the sample with a specified tag (barcode).
[0057] Specifically, the library of each sample contains connected inserts and tags corresponding specifically to the sample, the sequencing data includes reads of multiple samples, and the reads are split according to the part of the tag in the reads to obtain the sequencing data of each sample. Thus, the reads are split according to the specific tag sequence of the sample corresponding to the readout, and the reads assigned to the specified sample are obtained.
[0058] After obtaining the reads assigned to the specified sample, step S103 is performed, and according to the classification result of the pre-set classifier for the reads assigned to the specified sample, the corresponding part of the reads belonging to the specified category is filtered out.
[0059] The process includes obtaining the value of the feature of the reads assigned to the specified sample, and querying the pre-set classifier according to the value of the feature to determine the classification result of the reads assigned to the specified sample. The pre-set classifier is a quantitative scheme that associates the feature with the probability of the reads belonging to the specified category. The feature referred to is a plurality of quantifiable features generated in the sequencing run, which are related to the signal for obtaining the reads assigned to the specified sample and / or assigning the reads to the specified sample.
[0060] The values of the features of the split assigned reads of the designated sample are input into a preset classifier, which can determine the probability of the corresponding reads being a mismatched read (reads assigned to the designated sample by mistake) or a correctly assigned read (reads correctly assigned to the designated sample), which can be used to represent the confidence of the split being correct, and the higher the probability, the less likely the reads are assigned to the designated sample due to label jumping. Therefore, the reads assigned to the designated sample can be screened based on the probability score, for example, filtering out those reads determined by the preset classifier to have a higher probability of being a mismatched read, i.e. reads that the preset classifier considers to have a higher probability of being assigned by mistake due to label jumping.
[0061] In short, the preset classifier is a quantitative scheme for associating the features of the reads with the probability of the classification of the reads. In this paper, the preset classifier is also sometimes referred to as a model or a function or a relationship. The model predicts the probability of the reads being a mismatched read by inputting the feature values of the reads, and optionally determines the category of the reads based on the probability; based on the predicted classification result, the reads are further screened or filtered, which can more accurately obtain the reads originating from the designated sample, and further more accurately detect and analyze the designated sample.
[0062] The so-called features are a plurality of quantifiable features generated in the sequencing run, which are related to the signals obtained for the assignment of the reads to the designated sample and / or the assignment of the reads to the designated sample.
[0063] The preset classifier can be established during the process of multiple sample analysis, such as in the process of S101 or S102 or before S103 based on training data, or it can be pre-trained to call the prediction result of the reads by the preset classifier for S103.
[0064] In some examples, the preset classifier is pre-trained based on a training data set to construct and store for calling in multiple sample analysis.
[0065] In some examples, determining or constructing the preset classifier includes the following steps:
[0066] The training data set is obtained, which is sequencing data of multiple samples whose read assignment results and true value assignment results are known, and the training data set includes a plurality of reads containing at least a part of the sequence to be sequenced of each sample and the specific label sequence carried by the sample, and the training data set contains at least a certain number of reads;
[0067] The reads are divided into two categories according to whether the assignment result and the true value assignment result are consistent, and each read in the two categories is quantitatively represented in an alternative form based on the features of the reads.
[0068] The number of each type of reads contained in each alternative form is counted to establish the correspondence between the alternative form and the read type.
[0069] The training data set is sequencing data of multiple samples in which the sample assignment results and the true value assignment results of the reads are known, i.e., the training data set is sequencing data of multiple samples in which the normal reads (reads correctly assigned to the true source sample, sometimes also referred to herein as non-string reads) and abnormal reads (reads incorrectly assigned to a sample other than the true source sample, sometimes also referred to herein as mismatched reads or hopping-reads or jumping reads or string reads) are known. The more data contained in the training data set, the more accurate the prediction of the preset classifier for the type of reads. Therefore, in the embodiments of the present application, the training data set contains at least a certain number of reads, such as 10,000 reads, 100,000 reads, or even more reads. The reads can be divided into two categories according to whether the assignment result and the true value assignment result are consistent, i.e., into normal reads and abnormal reads. Then each read in the two types of reads is quantitatively represented as an alternative form according to the characteristics of the normal reads and the abnormal reads, thereby establishing the correspondence between the alternative form and the read type, so that the type of the read can be determined according to the alternative form when the prediction classifier is applied subsequently. Further, in order to accurately determine the type of the read according to the alternative form when the prediction classifier is applied subsequently, the number of each type of reads contained in each alternative form can also be counted to establish the correspondence between the alternative form and the read type.
[0070] In some examples, the training data set and the unsplit sequencing data in the multiple sample analysis method in any example can be obtained using the same or different sequencing methods, sequencing platforms and / or sequencing modes. The sequencing methods include but are not limited to sequencing by synthesis, sequencing by ligation or pyrosequencing; the sequencing platforms include but are not limited to the Illunima / solexa platform, the ABI Solid platform and the Roche 454 platform; the sequencing modes can use common sequencing modes of commercially available sequencing platforms such as SE-double index (single-end double index), SE-single index (single-end single index), PE-single index (double-end single index), PE-double index (double-end double index) and the like. Preferably, the training data set and the unsplit sequencing data in the multiple sample analysis method in any example are obtained from the same sequencing method or the same manufacturer's sequencing platform, which facilitates more accurate multiple sample analysis. Further, sequencing the multiple samples on different sequencers of the same sequencing platform can ensure that the training data set obtained does not have excessive bias.
[0071] The inventors generate features of tens of reads based on the information of images in the sequencing process, the identified bases, the sequence characteristics of the reads, etc. After screening and testing, some features are determined which can quantitatively represent the reads and have obvious differences in the distribution in the two types of reads, and are used to construct the preset classifier.
[0072] In some examples, the features of the reads are selected from at least one of the following a)-b): a) the signal intensity of the insert portion readout or any feature having a Spearman correlation coefficient with the signal intensity of the insert portion readout not less than 0.5; b) the signal intensity of the tag readout or any feature having a Spearman correlation coefficient with the signal intensity of the tag readout not less than 0.5.
[0073] In some examples, the features of the reads include a) and b); more specifically, the features of the reads include the signal intensity of the insert portion readout (hereinafter sometimes represented as readInts) and the signal intensity of the tag readout (hereinafter sometimes represented as barInts). The value of the feature of the signal intensity of the insert portion readout readInts can be represented by the average or median of the signal intensity of all or part of the bases of the insert portion readout on the corresponding image. The value of the feature of the signal intensity of the tag readout barInts can be represented by the average or median of the signal intensity of all the bases of the tag readout on the corresponding image.
[0074] The determination of the position of the extended or readout base on the corresponding image can be determined by using conventional methods such as the center of gravity method, etc. The signal intensity of the extended or readout base on the image can be absolute signal intensity or relative signal intensity. If not specifically stated, it generally refers to relative signal intensity. For example, the signal intensity after image preprocessing such as background removal (if it is a grayscale image, it can be equivalent to the grayscale value or pixel at the specified position, and if it is a color image, it can be converted into a grayscale image for calculation in order to quickly determine), such as the signal intensity after phasing, pre-phasing and / or crosstalk correction, etc. The signal intensity at the corresponding position can be determined by using the interpolation method. In the sequencing process based on the detection and identification of the extended base by fluorescence signal, the signal intensity is also sometimes referred to as brightness. In the sequencing platform based on surface fluorescence microscopic imaging detection, the signal intensity of a position corresponding to a chemical feature on the image collected (such as the bright spot feature presented on the image) in each round of reaction for the edge synthesis sequencing can be represented as an array of four bases, for example, represented as [B1, B2, B3, B4], wherein B1 to B4 are respectively the signals generated on A, T, C and G. For example, the signal intensity of a certain chemical feature position obtained in a round of reaction is [80, 10, 5, 5], and it is determined that the extended base type corresponding to the chemical feature position in this round of reaction is A, wherein 80 is the maximum signal intensity (or maximum brightness value; or maximum grayscale value; or maximum Ints), and 10 is the second largest signal intensity (or second largest brightness value; or second largest grayscale value; or second largest Ints). In a certain example, the insert portion readout is not less than 50 bp, and the maximum signal intensity of 16 to 25 bases on the corresponding image from the relative middle, for example, the 15th base, is sorted, and the median value represents the value of the readInts; the tag readout is not less than 6 bp, and the maximum signal intensity of each base on the corresponding image of the tag readout is sorted, and the median value represents the value of the barInts. In this way, the whole readout situation can be better reflected, and the value of the feature can also be quickly obtained. In some examples, the features of the read segment in addition to a) and b) also include the signal intensity of the specific base of the tag readout or any feature with a Spearman correlation coefficient with the signal intensity of the specific base of the tag readout not less than 0.5.In certain examples, the signal intensity of a particular base of a tag readout includes the maximum signal intensity of the particular base of the tag readout and / or the second largest signal intensity of the particular base of the tag readout, i.e., in addition to a), b), the features of the read segment also include one or more of c) and d): c) the maximum signal intensity of the particular base of the tag readout or any feature having a Spearman correlation coefficient with the maximum signal intensity of the particular base of the tag readout that is not less than 0.5; d) the second largest signal intensity of the particular base of the tag readout or any feature having a Spearman correlation coefficient with the second largest signal intensity of the particular base of the tag readout that is not less than 0.5. The maximum signal intensity of the particular base is sometimes denoted as badQMInts; the second largest signal intensity of the particular base is sometimes denoted as badQSInts. The particular base is, for example, the first base of the inverse Q value, the second base of the inverse Q value, or the third base of the inverse Q value.
[0075] In certain examples, the maximum signal intensity of a particular base of a tag readout or any feature having a Spearman correlation coefficient with the maximum signal intensity of the particular base of the tag readout that is not less than 0.5; and the second largest signal intensity of the particular base of the tag readout or any feature having a Spearman correlation coefficient with the second largest signal intensity of the particular base of the tag readout that is not less than 0.5 include:
[0076] the maximum signal intensity of the first base of the inverse Q value of a tag readout or any feature having a Spearman correlation coefficient with the maximum signal intensity of the first base of the inverse Q value of the tag readout that is not less than 0.5; and,
[0077] the second largest signal intensity of the first base of the inverse Q value of a tag readout or any feature having a Spearman correlation coefficient with the second largest signal intensity of the first base of the inverse Q value of the tag readout that is not less than 0.5.
[0078] In certain examples, the maximum signal intensity of a particular base of a tag readout or any feature having a Spearman correlation coefficient with the maximum signal intensity of the particular base of the tag readout that is not less than 0.5; and the second largest signal intensity of the particular base of the tag readout or any feature having a Spearman correlation coefficient with the second largest signal intensity of the particular base of the tag readout that is not less than 0.5 include:
[0079] the maximum signal intensity of the second base of the inverse Q value of a tag readout or any feature having a Spearman correlation coefficient with the maximum signal intensity of the second base of the inverse Q value of the tag readout that is not less than 0.5; and,
[0080] the second largest signal intensity of the second base of the inverse Q value of a tag readout or any feature having a Spearman correlation coefficient with the second largest signal intensity of the second base of the inverse Q value of the tag readout that is not less than 0.5.
[0081] In some examples, the maximum signal intensity of a specific base read by a tag or any feature with a Spearman correlation coefficient with the maximum signal intensity of the specific base read by the tag not less than 0.5; and the second maximum signal intensity of the specific base read by the tag or any feature with a Spearman correlation coefficient with the second maximum signal intensity of the specific base read by the tag not less than 0.5 include:
[0082] the maximum signal intensity of the third base from the bottom in terms of Q value read by a tag or any feature with a Spearman correlation coefficient with the maximum signal intensity of the third base from the bottom in terms of Q value read by the tag not less than 0.5; and
[0083] the second maximum signal intensity of the third base from the bottom in terms of Q value read by a tag or any feature with a Spearman correlation coefficient with the second maximum signal intensity of the third base from the bottom in terms of Q value read by the tag not less than 0.5.
[0084] In some examples, a quality score is assigned to each read base on a tag, i.e., a base quality value (Q value) is calculated for each read base to determine the reliability of each read base. The Q values of the read bases are sorted in descending order, and the maximum signal intensity and the second maximum signal intensity of the first, second or third base from the bottom in terms of Q value are taken as the feature values of badQMInts and badQSInts of the tag read, respectively. In this way, the overall read can be better reflected, and the feature values can be quickly obtained.
[0085] In some examples, the maximum signal intensity and the second maximum signal intensity of the second base from the bottom in terms of Q value are more stable than the maximum signal intensity and the second maximum signal intensity of the first base from the bottom in terms of Q value or the maximum signal intensity and the second maximum signal intensity of the third base from the bottom in terms of Q value, which can enhance the universality and practicality of the trained classifier.
[0086] In some examples, the feature values readInts, barInts, badQMInts and badQSInts are obtained by extracting the gray value or brightness value of the bright spot on the image based on surface fluorescence microscopic imaging detection during sequencing, and can be selected based on actual application requirements from the above predetermined features. Typically, a combination of the two feature values readInts and barInts can be selected. In some examples, in the case of a short tag sequence and high sequence similarity between tags, a combination of the four feature values readInts, barInts, badQMInts and badQSInts can be selected.
[0087] In some embodiments, quantifying each read in the two types of reads into an alternative form based on the features of the reads refers to a form identified by the features corresponding to the features of the processed reads, such as a normalized feature value, or a feature value revalued based on a weight coefficient corresponding to the feature value, and the like. Taking the normalized representation as an example, quantifying each read in the two types of reads into an alternative form based on the features of the reads includes: performing normalization processing on the feature values corresponding to the features of the reads to obtain normalized feature values; and quantifying the normalized feature values into an alternative form.
[0088] In some examples, quantifying each read in the two types of reads into an alternative form according to the features of the normal reads and the abnormal reads includes:
[0089] performing normalization processing on the feature values corresponding to the features of the normal reads and the abnormal reads to obtain normalized feature values, and quantifying the normalized feature values into an alternative form.
[0090] Specifically, performing normalization processing on the four feature values readInts, barInts, badQMInts, and badQSInts of each normal read and abnormal read to form four-dimensional discrete data from [0, 0, 0, 0] to [9, 9, 9, 9], and representing the discrete data formed after the normalization processing into an alternative form; or,
[0091] performing normalization processing on the two feature values readInts and barInts of each normal read and abnormal read to form two-dimensional discrete data from [0, 0] to [9, 9], and representing the discrete data formed after the normalization processing into an alternative form.
[0092] In some examples, the preset classifier constructed is shown in Table 1:
[0093] Table 1
[0094] wherein Type represents the discrete data formed after normalization of the feature values, Type is represented by Ai, i is the number of rows of the current table, for example, Ai can be 2239.
[0095] HoppedNum represents the number of abnormal reads in the sequencing data, for example, in Table 1, it is represented by Bi, i is the number of rows of the current table, for example, Bi can be 0, i.e., the number of abnormal reads is 0.
[0096] TotalNum represents the total number of reads in the sequencing data, in Table 1, it is represented by Ci, i is the number of rows of the current table, for example, Ci can be 29.
[0097] The score represents a probability, and the value is the integer part of -log(hoppedNum / totalNum), and the value range is [0, 9], which is used to represent the confidence of correct splitting. The higher the probability is, the smaller the possibility that the read is incorrectly allocated to the specified sample due to label jumping is. Conversely, the lower the probability is, the greater the possibility that the read is incorrectly allocated to the specified sample due to label jumping is.
[0098] In actual application, after a plurality of sample data are counted, a classification table of label jumping reads is generated, and a preset classifier can be formed based on data in the classification table.
[0099] After the read that helps the specified category is filtered out from the read allocated to the specified sample according to the preset classifier, the specified sample is detected according to the remaining read allocated to the specified sample according to step S104.
[0100] The remaining read is the read that is left after the read that is predicted by the preset classifier to be incorrectly allocated to the specified sample with a high probability is removed, and then the specified sample is detected according to the remaining read, to obtain biological information of the specified sample, for example, the read after screening is used to detect mutations of the specified sample or the organism from which the sample is derived. In this way, the influence of incorrect splitting of reads caused by label jumping can be reduced to a certain extent, and the accuracy of detecting the specified sample based on the read allocated to the specified sample is improved. Specifically, the biological information of the specified sample can be obtained by detecting the specified sample according to the remaining read allocated to the specified sample, for example, the base sequence order and base mutation information can be obtained.
[0101] In the embodiments of the present application, a multiple sample analysis device is also provided, as shown in FIG. 2, which comprises:
[0102] The acquisition unit 201 is configured to acquire sequencing data of multiple samples, and the sequencing data comprises reads from multiple samples.
[0103] The splitting unit 202 is configured to split the sequencing data to obtain reads allocated to a specified sample.
[0104] The filtering unit 203 is configured to filter out corresponding part of reads belonging to a specified category from the reads allocated to the specified sample according to a classification result of the reads allocated to the specified sample by a preset classifier, which comprises obtaining a value of a feature of the reads allocated to the specified sample, and querying the preset classifier according to the value of the feature to determine the classification result of the reads allocated to the specified sample.
[0105] wherein the pre-set classifier is a quantification scheme that associates the features with probabilities of the reads belonging to a specified class of reads, and the features are a plurality of quantifiable features generated in a sequencing run that are related to the reads assigned to the specified sample and / or to the signal that assigned the reads to the specified sample;
[0106] and a detecting unit 204 configured to detect the specified sample based on the remaining reads assigned to the specified sample.
[0107] In some examples, the obtaining unit comprises:
[0108] a constructing sub-unit configured to construct a multiplex library comprising a plurality of samples, the multiplex library being a mixture of sequencing libraries of the plurality of samples, each sequencing library of a sample being tagged with a tag specific to the sample, the tag being a known sequence of greater than or equal to 1 bp, and the tags specific to any two samples being different by at least one base; and
[0109] a sequencing sub-unit configured to sequence the multiplex library.
[0110] In some examples, the sequencing data further comprises sequencing sequences of the tags specific to the specified sample, and the sequencing data is split based on the sequencing sequences of the tags to obtain the reads assigned to the specified sample.
[0111] In some examples, the multiplex library comprises the tags and the inserts connected to the tags, and the obtaining unit is specifically configured to sequence the multiplex library based on surface fluorescence microscopic imaging detection and sequencing-by-synthesis, and the obtaining unit comprises:
[0112] an amplifying unit configured to connect the multiplex library to a surface and amplify the multiplex library on the surface, or amplify the multiplex library and load the amplified multiplex library to the surface, to form a surface with a plurality of amplicons;
[0113] a plurality of rounds of sequencing sub-units configured to place the surface in a condition suitable for polymerization reaction, sequence the amplicons in a plurality of rounds, including collecting images of each round of sequencing and detecting signals of positions of the amplicons on the images, to identify types of bases incorporated into a specified amplicon in the round of reaction, to determine sequencing sequences of at least part of the inserts and sequencing sequences of the tags in the specified amplicon.
[0114] In some examples, the features are selected from at least one of the following a)-d):
[0115] a) signal intensity of the sequencing sequence of the insert or any feature with a Spearman correlation coefficient with the signal intensity of the sequencing sequence of the insert not less than 0.5;
[0116] b) the signal intensity of the sequencing sequence of the tag or any feature with a Spearman correlation coefficient with the signal intensity of the sequencing sequence of the tag no less than 0.5;
[0117] c) the maximum signal intensity of a specific base of the tag read or any feature with a Spearman correlation coefficient with the maximum signal intensity of a specific base of the tag read no less than 0.5;
[0118] d) the second maximum signal intensity of a specific base of the tag read or any feature with a Spearman correlation coefficient with the second maximum signal intensity of a specific base of the tag read no less than 0.5.
[0119] In some examples, the apparatus further comprises a classifier determining unit, which comprises:
[0120] a data obtaining sub-unit configured to obtain a training data set, the training data set being sequencing data of multiple samples whose read and the sample from which the read originates have known assignment results and true value assignment results, the training data set comprising a plurality of reads comprising at least a part of a sequence to be sequenced of each sample and a specific tag sequence carried by the sample, the training data set comprising at least a specific number of reads;
[0121] a representation sub-unit configured to divide the reads into two categories according to whether the assignment results and the true value assignment results are consistent, and to represent each read in each category in an alternative form based on a feature of the read;
[0122] a establishing sub-unit configured to count the number of reads in each category contained in each alternative form, to establish a correspondence between the alternative form and the category of the read.
[0123] In some examples, the feature of the read comprises at least one of:
[0124] the signal intensity of the insert portion read or any feature with a Spearman correlation coefficient with the signal intensity of the insert read no less than 0.5;
[0125] the signal intensity of the tag read or any feature with a Spearman correlation coefficient with the signal intensity of the tag read no less than 0.5;
[0126] the maximum signal intensity of a specific base of the tag read or any feature with a Spearman correlation coefficient with the maximum signal intensity of a specific base of the tag read no less than 0.5;
[0127] the second maximum signal intensity of a specific base of the tag read or any feature with a Spearman correlation coefficient with the second maximum signal intensity of a specific base of the tag read no less than 0.5.
[0128] In certain examples, the maximum signal intensity for a particular base in the tag readout or any feature having a Spearman correlation coefficient with the maximum signal intensity for a particular base in the tag readout that is not less than 0.5; and,
[0129] the second largest signal intensity for a particular base in the tag readout or any feature having a Spearman correlation coefficient with the second largest signal intensity for a particular base in the tag readout that is not less than 0.5, includes:
[0130] the maximum signal intensity for the first base in the tag readout having the first inverse Q value or any feature having a Spearman correlation coefficient with the maximum signal intensity for the first base in the tag readout having the first inverse Q value that is not less than 0.5; and,
[0131] the second largest signal intensity for the first base in the tag readout having the first inverse Q value or any feature having a Spearman correlation coefficient with the second largest signal intensity for the first base in the tag readout having the first inverse Q value that is not less than 0.5.
[0132] In certain examples, the maximum signal intensity for a particular base in the tag readout or any feature having a Spearman correlation coefficient with the maximum signal intensity for a particular base in the tag readout that is not less than 0.5; and,
[0133] the second largest signal intensity for a particular base in the tag readout or any feature having a Spearman correlation coefficient with the second largest signal intensity for a particular base in the tag readout that is not less than 0.5, includes:
[0134] the maximum signal intensity for the second base in the tag readout having the second inverse Q value or any feature having a Spearman correlation coefficient with the maximum signal intensity for the second base in the tag readout having the second inverse Q value that is not less than 0.5; and,
[0135] the second largest signal intensity for the second base in the tag readout having the second inverse Q value or any feature having a Spearman correlation coefficient with the second largest signal intensity for the second base in the tag readout having the second inverse Q value that is not less than 0.5.
[0136] In certain examples, the maximum signal intensity for a particular base in the tag readout or any feature having a Spearman correlation coefficient with the maximum signal intensity for a particular base in the tag readout that is not less than 0.5; and,
[0137] the second largest signal intensity for a particular base in the tag readout or any feature having a Spearman correlation coefficient with the second largest signal intensity for a particular base in the tag readout that is not less than 0.5, includes:
[0138] the maximum signal intensity for the third base in the tag readout having the third inverse Q value or any feature having a Spearman correlation coefficient with the maximum signal intensity for the third base in the tag readout having the third inverse Q value that is not less than 0.5; and,
[0139] The second-smallest signal intensity of the base whose Q value is the third from the largest in the tag readout or a Spearman correlation coefficient with the second-smallest signal intensity of the base whose Q value is the third from the largest in the tag readout is not less than 0.5.
[0140] In some examples, the representation subunit is configured to:
[0141] The feature value corresponding to the feature of the read is normalized to obtain a normalized feature value.
[0142] The normalized feature value is quantitatively represented in an alternative form.
[0143] It should be noted that the specific implementation of each unit and subunit in this embodiment can refer to the corresponding content in the foregoing, which will not be described in detail here.
[0144] In another embodiment of the present application, a computing device is also provided, which comprises:
[0145] A memory for storing a computer executable program;
[0146] One or more processors for executing the computer executable program and implementing the multiple sample analysis method according to any one of the above.
[0147] In another embodiment of the present application, a computer readable storage medium for storing a computer executable program is also provided, and executing the program comprises completing the multiple sample analysis method according to any one of the above.
[0148] It should be noted that the specific implementation of the processor in this embodiment can refer to the corresponding content in the foregoing, which will not be described in detail here.
[0149] In still another embodiment of the present application, a method for constructing a classifier capable of predicting whether a read of mixed sample sequencing data split to a specified sample is a string sample read or a non-string sample read is also provided, and the method for constructing the classifier comprises:
[0150] Obtaining a training data set, which is sequencing data of multiple samples whose read sample assignments are known and whose true value assignments are known, the training data set comprising a plurality of reads comprising at least a portion of a sequence to be sequenced of each sample and a specific tag sequence carried by the sample, and the training data set comprising at least a specific number of reads;
[0151] According to whether the assignment result and the true value assignment result are consistent, the reads are divided into two categories, and each read in the two categories of reads is quantitatively represented in an alternative form based on the features of the reads.
[0152] counting the number of each type of reads contained in each alternative form to establish the correspondence between the alternative form and the read type;
[0153] based on the correspondence, constructing a classifier.
[0154] The training data set is known in which the reads and the samples from which they originate, i.e. the training data set is a sample with known labels, even if the label jump is generated, it can be identified. The more the number of training data sets, the higher the accuracy of the analysis based on the statistical results, therefore, in the embodiments of the present application, the training data set contains at least a certain number of reads, for example, 10,000 reads, 100,000 reads or even more reads. The true value assignment result represents the reads that generate label jumps and the reads that do not generate label jumps. Whether the assignment result and the true value assignment result are consistent will divide the reads into two categories, i.e. into reads that generate label jumps and reads that do not generate label jumps. Then according to the characteristics of the reads, each read in the two categories of reads is quantitatively represented as an alternative form, so that the type of the read can be determined according to the alternative form, and further in order to enable the alternative form to be accurately applied in subsequent read type identification, the number of each type of reads contained in each alternative form can be counted to establish the correspondence between the alternative form and the read type.
[0155] In some examples, the characteristics of the reads include at least one of the following:
[0156] any feature of which the Spearman correlation coefficient of the signal intensity of the insert segment part readout or the signal intensity of the insert segment readout is not less than 0.5;
[0157] any feature of which the signal intensity of the tag readout or the Spearman correlation coefficient of the signal intensity of the tag readout is not less than 0.5;
[0158] any feature of which the maximum signal intensity of the specific base of the tag readout or the Spearman correlation coefficient of the maximum signal intensity of the specific base of the tag readout is not less than 0.5;
[0159] any feature of which the second largest signal intensity of the specific base of the tag readout or the Spearman correlation coefficient of the second largest signal intensity of the specific base of the tag readout is not less than 0.5.
[0160] In some examples, the maximum signal intensity of the specific base of the tag readout or the Spearman correlation coefficient of the maximum signal intensity of the specific base of the tag readout is not less than 0.5; and,
[0161] the second largest signal intensity of the specific base of the tag readout or the Spearman correlation coefficient of the second largest signal intensity of the specific base of the tag readout is not less than 0.5, including:
[0162] the maximum signal intensity of the base with the first smallest reciprocal Q value read by the tag or any feature with a Spearman correlation coefficient with the maximum signal intensity of the base with the first smallest reciprocal Q value read by the tag not less than 0.5; and
[0163] the second largest signal intensity of the base with the first smallest reciprocal Q value read by the tag or any feature with a Spearman correlation coefficient with the second largest signal intensity of the base with the first smallest reciprocal Q value read by the tag not less than 0.5.
[0164] In certain examples, the maximum signal intensity of a particular base read by the tag or any feature with a Spearman correlation coefficient with the maximum signal intensity of the particular base read by the tag not less than 0.5; and
[0165] the second largest signal intensity of a particular base read by the tag or any feature with a Spearman correlation coefficient with the second largest signal intensity of the particular base read by the tag not less than 0.5, including:
[0166] the maximum signal intensity of the base with the second smallest reciprocal Q value read by the tag or any feature with a Spearman correlation coefficient with the maximum signal intensity of the base with the second smallest reciprocal Q value read by the tag not less than 0.5; and
[0167] the second largest signal intensity of the base with the second smallest reciprocal Q value read by the tag or any feature with a Spearman correlation coefficient with the second largest signal intensity of the base with the second smallest reciprocal Q value read by the tag not less than 0.5.
[0168] In certain examples, the maximum signal intensity of a particular base read by the tag or any feature with a Spearman correlation coefficient with the maximum signal intensity of the particular base read by the tag not less than 0.5; and
[0169] the second largest signal intensity of a particular base read by the tag or any feature with a Spearman correlation coefficient with the second largest signal intensity of the particular base read by the tag not less than 0.5, including:
[0170] the maximum signal intensity of the base with the third smallest reciprocal Q value read by the tag or any feature with a Spearman correlation coefficient with the maximum signal intensity of the base with the third smallest reciprocal Q value read by the tag not less than 0.5; and
[0171] the second largest signal intensity of the base with the third smallest reciprocal Q value read by the tag or any feature with a Spearman correlation coefficient with the second largest signal intensity of the base with the third smallest reciprocal Q value read by the tag not less than 0.5.
[0172] In some examples, in order to make the data quantization more accurate, the alternative form can be determined based on normalization. In the embodiments of the present application, and based on the predetermined features, each read in the two categories is quantized into an alternative form, including: normalizing the feature values corresponding to the predetermined features to obtain normalized feature values; and quantizing the normalized feature values into an alternative form.
[0173] In another embodiment of the present application, a computer readable storage medium is also provided for storing a program for computer execution, and execution of the program includes completing the method of constructing a classifier as described above.
[0174] In still another embodiment of the present application, a computer device is also provided, which includes:
[0175] a memory for storing a computer executable program;
[0176] one or more processors for executing the computer executable program and implementing the method of constructing a classifier as described above.
[0177] In another embodiment of the present application, a method for applying a classifier is also provided, which includes:
[0178] classifying the reads assigned to the specified sample to obtain a classification result;
[0179] so that according to the classification result, the corresponding part of reads belonging to the specified category is filtered out, including obtaining the values of the features of the reads assigned to the specified sample, and querying the pre-set classifier according to the values of the features to determine the classification result of the reads assigned to the specified sample; and detecting the specified sample according to the remaining reads assigned to the specified sample;
[0180] wherein the reads assigned to the specified sample are obtained by splitting the sequencing data of the multiple samples, and the sequencing data includes reads from multiple samples.
[0181] The multiple sample analysis method or device disclosed in the embodiments of the present application obtains the classification result of the reads assigned to the specified sample through the pre-set classifier and filters out the part of reads belonging to the specified category, for example, the part of reads that are incorrectly assigned to the specified sample due to label jumping with a larger probability, thereby reducing the influence of label jumping and incorrectly assigned reads on the detection of the specified sample.
[0182] The various embodiments described in this specification are intended to be exemplary only. The subject matter described in this specification can be implemented in software, hardware, or a combination thereof. The order in which the methods are described is not intended to be a limitation unless otherwise specified. Any procedures, features, steps, or blocks described herein can be performed in an order other than as described unless otherwise specified or implied. Any procedures, features, steps, or blocks described herein can be combined with each other, other procedures, features, steps, or blocks unless otherwise specified or implied.
[0183] Those skilled in the art will further appreciate that the units and algorithms described in connection with the examples disclosed herein can be embodied directly in hardware, software, or a combination thereof. To the extent that they are implemented in software, the functions described can be stored in or implemented as one or more software modules in RAM, memory, mass storage, or the like, for execution by a processor. The software modules can be executed by a processor of a general purpose computer, a special purpose computer, or a network computer. To the extent that they are implemented in hardware, the functions described can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, micro-controllers, microprocessors, or other hardware devices suitable for implementation.
[0184] The processes and algorithms described in connection with the examples disclosed herein can be embodied directly in hardware, software, or a combination thereof. To the extent that they are implemented in software, the functions described can be stored in or implemented as one or more software modules in RAM, memory, mass storage, or the like, for execution by a processor. The software modules can be executed by a processor of a general purpose computer, a special purpose computer, or a network computer. To the extent that they are implemented in hardware, the functions described can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, micro-controllers, microprocessors, or other hardware devices suitable for implementation.
[0185] The above description is intended to enable any person skilled in the art to make or use the application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multiplex sample analysis method, characterized by, The method comprises: obtaining sequencing data of multiple samples, the sequencing data comprising reads from multiple samples; splitting the sequencing data to obtain reads assigned to a specified sample; filtering out corresponding part of reads belonging to a specified category from the assigned reads to the specified sample according to a classification result of the assigned reads to the specified sample by a preset classifier, comprising obtaining values of features of the assigned reads to the specified sample, and querying the preset classifier according to the values of the features to determine the classification result of the assigned reads to the specified sample, wherein the preset classifier is a quantization scheme associated with the features and the probability of the assigned reads to the specified sample belonging to the specified category, and the features are multiple quantifiable features generated in a sequencing run and related to signals for obtaining the assigned reads to the specified sample and / or assigning the reads to the specified sample; and detecting the specified sample according to the remaining assigned reads to the specified sample.
2. The method of claim 1, wherein, The method of obtaining sequencing data of multiple samples comprises: constructing a multiplex library comprising multiple samples, the multiplex library being a mixture of sequencing libraries of multiple samples, each sequencing library of a sample being labeled with a label specific to the sample, the label being a known sequence of greater than or equal to 1 bp, and the labels specific to any two samples having a difference of at least one base; and sequencing the multiplex library.
3. The method according to claim 1 or 2, characterized in that, The sequencing data further comprises sequencing sequences of the labels specific to the specified sample, and the sequencing data is split according to the sequencing sequences of the labels to obtain reads assigned to the specified sample.
4. The method of claim 3, wherein, The multiplex library comprises connected inserts and labels, and the multiplex library is sequenced by sequencing-by-synthesis based on surface fluorescence microscopic imaging detection, comprising: connecting the multiplex library to a surface and amplifying the multiplex library on the surface, or amplifying the multiplex library and loading the amplified multiplex library to the surface, to form a surface with multiple amplicons; placing the surface under conditions suitable for polymerization reaction, and performing multiple rounds of sequencing on the amplicons, comprising collecting images of each round of sequencing and detecting signals of positions of the corresponding amplicons on the images to identify types of bases incorporated into a specified amplicon in the round of reaction, to determine sequencing sequences of at least part of the inserts and sequencing sequences of the labels in the specified amplicon.
5. The method of claim 4, wherein, The features are selected from at least one of the following a)-d): a) signal intensity of the sequencing sequence of the insert or any feature with a Spearman correlation coefficient with the signal intensity of the sequencing sequence of the insert not less than 0.5; b) signal intensity of the sequencing sequence of the label or any feature with a Spearman correlation coefficient with the signal intensity of the sequencing sequence of the label not less than 0.5; c) maximum signal intensity of a specific base read by the label or any feature with a Spearman correlation coefficient with the maximum signal intensity of the specific base read by the label not less than 0.5; d) second maximum signal intensity of a specific base read by the label or any feature with a Spearman correlation coefficient with the second maximum signal intensity of the specific base read by the label not less than 0.
5.
6. The method according to any one of claims 1 to 5, characterized in that, The preset classifier is determined by the following steps: Obtaining a training data set, the training data set being sequencing data of multiple samples of reads of which sample assignments are known and true value assignments, the training data set comprising a plurality of reads comprising at least a portion of a sequence to be sequenced of each sample and a specific tag sequence carried by the sample, the training data set comprising at least a specific number of reads; According to whether the assignment result and the true value assignment result are consistent, the reads are divided into two categories, and each read in the two categories of reads is quantitatively represented in an alternative form based on the characteristics of the reads; The number of each category of reads contained in each alternative form is counted to establish the correspondence between the alternative form and the read category.
7. The method of claim 6, wherein, The characteristics of the reads include at least one of the following: The signal intensity of the insert portion read or the Spearman correlation coefficient of the signal intensity of the insert read is not less than 0.5; The signal intensity of the tag read or the Spearman correlation coefficient of the signal intensity of the tag read is not less than 0.5; The maximum signal intensity of a specific base of the tag read or the Spearman correlation coefficient of the maximum signal intensity of the specific base of the tag read is not less than 0.5; The second maximum signal intensity of a specific base of the tag read or the Spearman correlation coefficient of the second maximum signal intensity of the specific base of the tag read is not less than 0.
5.
8. The method of claim 7, wherein, The maximum signal intensity of a specific base of the tag read or the Spearman correlation coefficient of the maximum signal intensity of the specific base of the tag read is not less than 0.5; and The second maximum signal intensity of a specific base of the tag read or the Spearman correlation coefficient of the second maximum signal intensity of the specific base of the tag read is not less than 0.
5. The maximum signal intensity of a specific base of the tag read or the Spearman correlation coefficient of the maximum signal intensity of the specific base of the tag read is not less than 0.5; and 9. The method of claim 7, wherein, The second maximum signal intensity of a specific base of the tag read or the Spearman correlation coefficient of the second maximum signal intensity of the specific base of the tag read is not less than 0.
5. The maximum signal intensity of a specific base of the tag read or the Spearman correlation coefficient of the maximum signal intensity of the specific base of the tag read is not less than 0.5; and The second maximum signal intensity of a specific base of the tag read or the Spearman correlation coefficient of the second maximum signal intensity of the specific base of the tag read is not less than 0.
5. The maximum signal intensity of a specific base of the tag read or the Spearman correlation coefficient of the maximum signal intensity of the specific base of the tag read is not less than 0.5; and The second maximum signal intensity of a specific base of the tag read or the Spearman correlation coefficient of the second maximum signal intensity of the specific base of the tag read is not less than 0.
5.
10. The method of claim 7, wherein, a maximum signal intensity of a specific base in the tag readout or any feature with a Spearman correlation coefficient with the maximum signal intensity of the specific base in the tag readout not less than 0.5; and a second maximum signal intensity of a specific base in the tag readout or any feature with a Spearman correlation coefficient with the second maximum signal intensity of the specific base in the tag readout not less than 0.5, including: a maximum signal intensity of a Q value third from the last base in the tag readout or any feature with a Spearman correlation coefficient with the maximum signal intensity of the Q value third from the last base in the tag readout not less than 0.5; and a second maximum signal intensity of a Q value third from the last base in the tag readout or any feature with a Spearman correlation coefficient with the second maximum signal intensity of the Q value third from the last base in the tag readout not less than 0.
5.
11. The method according to any one of claims 7-10, characterized in that, quantifying and representing each read in the two categories of reads into an alternative form based on the read-based features includes: normalizing the feature values corresponding to the read-based features to obtain normalized feature values; quantifying and representing the normalized feature values into an alternative form.
12. A multiplexed sample analysis device, characterized by, including: an acquisition unit configured to acquire sequencing data of a multiplex sample, the sequencing data including reads from a plurality of samples; a splitting unit configured to split the sequencing data to obtain reads assigned to a specified sample; a filtering unit configured to filter out a corresponding part of reads belonging to a specified category of reads from the reads assigned to the specified sample according to a classification result of the reads assigned to the specified sample by a preset classifier, including acquiring a value of a feature of the reads assigned to the specified sample, and querying the preset classifier according to the value of the feature to determine the classification result of the reads assigned to the specified sample, wherein the preset classifier is a quantization scheme associated with the feature and a probability of the reads assigned to the specified sample belonging to the specified category of reads, and the feature is a plurality of quantifiable features generated in a sequencing run and related to signals for acquiring the reads assigned to the specified sample and / or assigning the reads to the specified sample; and a detection unit configured to detect the specified sample according to the remaining reads assigned to the specified sample.
13. The apparatus of claim 12, wherein, The acquisition unit includes: a construction subunit configured to construct a multiplex library containing a plurality of samples, the multiplex library being a mixture of sequencing libraries of the plurality of samples, each sequencing library of a sample being labeled with a tag specific to the sample, the tag being a known sequence of greater than or equal to 1 bp, and the tags specific to any two samples having a difference of at least one base; and a sequencing subunit configured to sequence the multiplex library.
14. The apparatus of claim 12 or 13, wherein, The sequencing data further includes a sequencing sequence of a tag specific to a specified sample, and the sequencing data is split according to the sequencing sequence of the tag to obtain reads assigned to the specified sample.
15. The apparatus of claim 14, wherein, The multiplex library contains connected inserts and the tags, and the acquisition unit is specifically configured to sequence the multiplex library by sequencing-by-synthesis based on surface fluorescence microscopic imaging detection, and the acquisition unit includes: an amplification unit configured to connect the multiplex library to a surface and amplify the multiplex library on the surface, or to amplify the multiplex library and load the amplified multiplex library to the surface, to form a surface with a plurality of amplicons; a multi-round sequencing subunit configured to place the surface under conditions suitable for polymerization reaction, to perform multi-round sequencing on the amplicons, including collecting images of each round of sequencing and detecting signals of positions of the corresponding amplicons on the images, to identify types of bases incorporated into a specified amplicon in the round of reaction, to determine a sequencing sequence of at least a portion of the insert fragment and a sequencing sequence of the tag in the specified amplicon.
16. The apparatus of claim 15, wherein, The features are selected from at least one of a) to d) below: a) a signal intensity of the sequencing sequence of the insert fragment or any feature with a Spearman correlation coefficient with the signal intensity of the sequencing sequence of the insert fragment no less than 0.5; b) a signal intensity of the sequencing sequence of the tag or any feature with a Spearman correlation coefficient with the signal intensity of the sequencing sequence of the tag no less than 0.5; c) a maximum signal intensity of a specific base of the tag read or any feature with a Spearman correlation coefficient with the maximum signal intensity of the specific base of the tag read no less than 0.5; d) a second maximum signal intensity of a specific base of the tag read or any feature with a Spearman correlation coefficient with the second maximum signal intensity of the specific base of the tag read no less than 0.
5.
17. The apparatus of any of claims 12-16, wherein, The apparatus further comprises a classifier determining unit, which comprises: a data obtaining subunit configured to obtain a training data set, the training data set being sequencing data of a plurality of samples with known allocation results of reads to samples of origins and true value allocation results, the training data set comprising a plurality of reads comprising at least a portion of a sequence to be sequenced of each sample and a specific tag sequence carried by the sample, the training data set comprising at least a specific number of reads; a representation subunit configured to divide the reads into two categories according to whether the allocation results and the true value allocation results are consistent, and to represent each read in each category in an alternative form based on features of the read; a establishing subunit configured to count the number of reads in each category contained in each alternative form, to establish a correspondence between the alternative form and the category of the read.
18. The apparatus of claim 17, wherein, The features of the read comprise at least one of: a signal intensity of a portion of the insert read or any feature with a Spearman correlation coefficient with the signal intensity of the insert read no less than 0.5; a signal intensity of the tag read or any feature with a Spearman correlation coefficient with the signal intensity of the tag read no less than 0.5; a maximum signal intensity of a specific base of the tag read or any feature with a Spearman correlation coefficient with the maximum signal intensity of the specific base of the tag read no less than 0.5; a second maximum signal intensity of a specific base of the tag read or any feature with a Spearman correlation coefficient with the second maximum signal intensity of the specific base of the tag read no less than 0.
5.
19. The method of claim 18, wherein, a maximum signal intensity of a specific base in the tag readout or any feature with a Spearman correlation coefficient with the maximum signal intensity of the specific base in the tag readout not less than 0.5; and a second largest signal intensity of the specific base in the tag readout or any feature with a Spearman correlation coefficient with the second largest signal intensity of the specific base in the tag readout not less than 0.5, including: a maximum signal intensity of a base with a first Q value reciprocal in the tag readout or any feature with a Spearman correlation coefficient with the maximum signal intensity of the base with the first Q value reciprocal in the tag readout not less than 0.5; and a second largest signal intensity of the base with the first Q value reciprocal in the tag readout or any feature with a Spearman correlation coefficient with the second largest signal intensity of the base with the first Q value reciprocal in the tag readout not less than 0.
5.
20. The method of claim 18, wherein a maximum signal intensity of a specific base in the tag readout or any feature with a Spearman correlation coefficient with the maximum signal intensity of the specific base in the tag readout not less than 0.5; and a second largest signal intensity of the specific base in the tag readout or any feature with a Spearman correlation coefficient with the second largest signal intensity of the specific base in the tag readout not less than 0.5, including: a maximum signal intensity of a base with a second Q value reciprocal in the tag readout or any feature with a Spearman correlation coefficient with the maximum signal intensity of the base with the second Q value reciprocal in the tag readout not less than 0.5; and a second largest signal intensity of the base with the second Q value reciprocal in the tag readout or any feature with a Spearman correlation coefficient with the second largest signal intensity of the base with the second Q value reciprocal in the tag readout not less than 0.
5.
21. The method of claim 18, wherein a maximum signal intensity of a specific base in the tag readout or any feature with a Spearman correlation coefficient with the maximum signal intensity of the specific base in the tag readout not less than 0.5; and a second largest signal intensity of the specific base in the tag readout or any feature with a Spearman correlation coefficient with the second largest signal intensity of the specific base in the tag readout not less than 0.5, including: a maximum signal intensity of a base with a third Q value reciprocal in the tag readout or any feature with a Spearman correlation coefficient with the maximum signal intensity of the base with the third Q value reciprocal in the tag readout not less than 0.5; and a second largest signal intensity of the base with the third Q value reciprocal in the tag readout or any feature with a Spearman correlation coefficient with the second largest signal intensity of the base with the third Q value reciprocal in the tag readout not less than 0.
5.
22. The apparatus of any of claims 18-21, wherein, The representation sub-unit is used to: normalize the feature value corresponding to the feature of the read, to obtain a normalized feature value; quantize the normalized feature value into an alternative form.
23. A computer readable storage medium for storing a program for execution by a computer, execution of the program comprising completing the multiple sample analysis method of any one of 1-11 above.
24. A computing device comprising: a memory for storing a computer executable program; a processor for executing the computer executable program; and a display for displaying the result of the execution of the computer executable program. one or more processors to execute the computer executable programs to implement the multiple sample analysis method as described in any of the above 1-11.
Citation Information
Patent Citations
High-throughput sequencing method based on internal reference of known tag
CN113981056A
Reference-guided genomic sequencing
CN114730617A
Method for determining read quality and sequencing method
CN117912550A
Multiplexed droplet-based sequencing using natural genetic barcodes
US20220005547A1