Multi-sample analysis method and device, storage medium and equipment

By constructing a pre-defined classifier, mismatched reads are predicted and filtered using read features, thus solving the problem of false positives and false negatives caused by label skipping in multiplex sequencing and improving the accuracy of sample detection.

CN120833852APending Publication Date: 2025-10-24GENEMIND BIOSCIENCES CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202410497278.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-19
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

The increase in false positives or false negatives due to tag skipping in multiplex sequencing, especially in applications such as cancer genomics and liquid biopsy that require accurate detection of rare variants, makes it difficult for existing technologies to effectively reduce tag assignment errors.

Method used

By constructing a pre-defined classifier, the features of the read segments in the training dataset are used to predict whether a read segment is a mismatched read segment, and the misassigned read segments are filtered out, thereby reducing the impact of label skipping on sample detection.

Benefits of technology

It effectively reduces misassignment caused by label skipping, improves the accuracy and reliability of sample detection, and reduces false positive and false negative results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833852A_ABST
    Figure CN120833852A_ABST
Patent Text Reader

Abstract

The invention discloses a multiple sample analysis method and device, a storage medium and equipment. The method comprises the following steps: acquiring sequencing data of multiple samples, wherein the sequencing data comprises reads from the multiple samples; splitting the sequencing data to obtain a read section distributed to a specified sample; and according to a classification result of the read segments allocated to the specified sample by a preset classifier, filtering out a part of read segments corresponding to the read segments belonging to the specified category, and detecting the specified sample according to the remaining read segments allocated to the specified sample. The preset classifier is a quantitative scheme associated with the probability that the features and the read belong to the read of the specified category. According to the method provided by the invention, the interference of the read segments (mismatched read segments) on subsequent detection and analysis is reduced by identifying the read segments which have high probability of tag jump and are more likely to be wrongly allocated to non-source samples and removing the read segments, so that the accuracy of multiple sample analysis and detection based on sequencing can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioinformatics, and more particularly, to a multiplex sample analysis method and device, a method of constructing a classifier capable of predicting read categories, a terminal or computing device, and a computer-readable storage medium. BACKGROUND

[0002] Next-generation sequencing, also known as high-throughput sequencing or massively parallel sequencing, can determine the nucleic acid sequences of multiple samples in one sequencing run. Such determination can be achieved through multiplex sample analysis, also known as multiplex library or multiplex sequencing. In multiplex sequencing, a specific sequence unique to the sample from which each nucleic acid fragment comes is added to each nucleic acid fragment during library construction, so that the libraries of multiple samples can be mixed in one reaction system for sequencing to obtain sequencing data, and then the sequencing data can be allocated to the corresponding sample according to the specific sequence, so as to obtain the sequencing data of each sample. The specific sequence is usually referred to as a tag or index (barcode or tag).

[0003] Tag misassignment between multiplex libraries, also known as index hopping or index misassignment or sample cross-talk, is a problem in multiplex sequencing. Understandably, any detection involving the use of high-throughput sequencing to seek trace amounts of "positive" data in a high-background-noise mixture is very susceptible to index hopping, including cancer genomics and other applications that require accurate detection of rare variants, such as liquid biopsy, pathogen detection, etc. For example, in pathogen detection applications, even if a single-digit number of reads of a certain pathogen is detected, it is generally considered as a positive detection of the pathogen. In this case, if the reads of one pathogen are mistakenly divided into another pathogen, i.e., the tag misassignment causes the reads from pathogen A to be mistakenly allocated to another pathogen B, this will lead to an increase in false positives or false negatives.

[0004] As can be seen, reducing index hopping and reducing misassignment are currently important issues in sequencing applications. SUMMARY

[0005] Therefore, embodiments of the present application provide a multiplex sample analysis method, a multiplex sample analysis device, a method of constructing a classifier capable of predicting read categories, a terminal or computing device, and a computer-readable storage medium.

[0006] The classifier construction method and multiple sample analysis method of the present application embodiment are based on the following conclusions and discoveries made by the inventors after observing and comparing a large amount of sequencing data with and without label hopping:

[0007] In common multiplexed sample analysis based on surface fluorescence imaging sequencing technology, a specific correspondence between index and sample is established before the libraries of each sample are mixed. Each library contains an index and an insert (a nucleic acid fragment of the sample to be tested, often referred to as an insert, test fragment, test sequence, insert, or insertion in library construction). Therefore, after obtaining sequencing data for the mixed sample, the associated insert read is usually assigned to the corresponding sample based on the index read to obtain sequencing data for the specified sample for detection and analysis. This process is also generally referred to as read splitting, allocation, or demultiplexing.

[0008] For a surface having multiple clusters to be sequenced (amplified products of a library), ideally or generally, a cluster includes multiple polynucleotide molecules with the same sequence, and sequencing the cluster can produce a group of reads with the same sequence (hereinafter sometimes referred to as the same read). By observing the conversion and acquisition process of the control signal, such as the conversion of the biochemical reaction signal into an image signal and the final process of determining the base and determining the read based on the image signal, the inventors summarized and found that if the image information after acquisition and processing, the signal changes and detection results of the designated cluster where the biochemical reaction occurred are shown on the image, compared with the normal detection signal and identification results of the same test part of a cluster, which are mostly "one cluster, one index, one read", the signal changes or identification results of the cluster corresponding to the reads that have been misassigned due to label jumping can be divided into the following four types: reading one index and one read from a cluster and the read quality of the index and / or reads is low, reading multiple indices and one read from a cluster, reading one index and multiple reads from a cluster, and reading multiple indices and multiple reads from a cluster. Of course, the detection signals and identification results of the test part of some clusters have multiple of the above situations at the same time. Moreover, after statistics, the inventors also found that the occurrence of any of the last three situations is likely to cause the sequencing data to be incorrectly split into samples that are not their true source.

[0009] Based on the summary findings, the inventors utilized training data with normal reads (reads that are correctly assigned to their true source sample, sometimes also referred to herein as non-string reads) and abnormal reads (reads that are incorrectly assigned to a sample other than their true source sample, sometimes also referred to herein as mismatched reads or hopping- or jumping- or stringed reads) as classification labels, constructed, tested and screened a plurality of features of the reads, particularly those that are highly associated with the observed cases described above, and ultimately determined the features of the reads that can better distinguish or fit the distribution of the two types of labels to train the function relationship / model of the classification labels and features.

[0010] The model was verified to be able to accurately predict the classification results of the reads (whether they are mismatched reads or the probability of being mismatched reads or normal reads). For this purpose, the embodiments of the present application provide a multiplex sample analysis method, comprising:

[0011] obtaining sequencing data of a multiplex sample, the sequencing data comprising reads from a plurality of samples;

[0012] splitting the sequencing data to obtain reads assigned to a specified sample;

[0013] filtering out a corresponding part of reads belonging to a specified category of reads according to the classification results of the reads assigned to the specified sample by a preset classifier, comprising obtaining the values of features of the reads assigned to the specified sample, and querying the preset classifier according to the values of the features to determine the classification results of the reads assigned to the specified sample,

[0014] wherein the preset classifier is a quantization scheme associated with the features and the probability of the reads belonging to the specified category of reads, and the features are a plurality of quantifiable features generated during sequencing runs and related to the signals for obtaining the reads assigned to the specified sample and / or assigning the reads to the specified sample;

[0015] and detecting the specified sample according to the remaining reads assigned to the specified sample.

[0016] A multiplex sample analysis device, comprising:

[0017] an obtaining unit configured to obtain sequencing data of a multiplex sample, the sequencing data comprising reads from a plurality of samples;

[0018] a splitting unit configured to split the sequencing data to obtain reads assigned to a specified sample;

[0019] a filtering unit configured to filter out a corresponding portion of reads belonging to a specified category of reads from the reads assigned to a specified sample according to a classification result of the reads assigned to the specified sample by a preset classifier, including obtaining a value of a feature of the reads assigned to the specified sample, and querying the preset classifier according to the value of the feature to determine the classification result of the reads assigned to the specified sample,

[0020] wherein the preset classifier is a quantization scheme associated with the feature and a probability of the reads belonging to the specified category of reads, and the feature is a plurality of quantifiable features generated in a sequencing run and related to a signal for obtaining the reads assigned to the specified sample and / or assigning the reads to the specified sample;

[0021] and a detecting unit configured to detect the specified sample according to the remaining reads assigned to the specified sample.

[0022] A computer readable storage medium for storing a program executable by a computer, and executing the program includes completing the multiple sample analysis method according to any one of the above.

[0023] A computing device, comprising:

[0024] a memory configured to store a computer executable program;

[0025] one or more processors configured to execute the computer executable program and implement the multiple sample analysis method according to any one of the above.

[0026] A method for constructing a classifier capable of predicting whether reads assigned to a specified sample from mixed sample sequencing data are string reads or non-string reads, the method comprising:

[0027] obtaining a training data set, the training data set being sequencing data of multiple samples with known sample assignment results of the reads and true value assignment results thereof, the training data set including a plurality of reads containing at least a portion of a sequence to be sequenced of each sample and a specific label sequence carried by the sample, and the training data set containing at least a certain number of reads;

[0028] dividing the reads into two categories according to whether the assignment results and the true value assignment results are consistent, and quantifying each read in the two categories of reads into an alternative form based on a feature of the read;

[0029] counting the number of each category of reads contained in each alternative form to establish a corresponding relationship between the alternative form and the category of reads;

[0030] constructing a classification classifier based on the corresponding relationship.

[0031] A computer readable storage medium for storing a program for execution by a computer, execution of the program comprising carrying out the method of building a classifier as described above.

[0032] A computer device comprising:

[0033] a memory for storing a computer executable program;

[0034] one or more processors for executing the computer executable program and implementing the method of building a classifier as described above.

[0035] A method of applying a classifier, comprising:

[0036] classifying reads assigned to a given sample to obtain a classification result;

[0037] such that, according to the classification result, a corresponding portion of reads belonging to a given class is filtered out, comprising obtaining a value of a feature of the reads assigned to the given sample, and querying the pre-set classifier according to the value of the feature to determine the classification result of the reads assigned to the given sample; and detecting the given sample according to the remaining reads assigned to the given sample;

[0038] wherein the reads assigned to the given sample are obtained by splitting sequencing data of multiple samples, and the sequencing data comprises reads from multiple samples.

[0039] The multiple sample analysis method or device disclosed in the embodiments of the present application obtains a classification result of reads assigned to a given sample by a pre-set classifier and filters out a portion of reads belonging to a given class, for example, those reads that are incorrectly assigned to the given sample due to a label jump with a relatively high probability, thereby reducing the impact of label jumps and incorrectly assigned reads on detection of the given sample. BRIEF DESCRIPTION OF DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort based on the provided drawings.

[0041] Figure 1 A flowchart of a multiple sample analysis method provided by an embodiment of the present application;

[0042] Figure 2 A structural diagram of a multiple sample analysis device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0043] The technical solutions in the embodiments of the present application will be described in detail below with reference to the drawings in the embodiments of the present application. The described embodiments are exemplary and are only used to explain the present application, and cannot be understood as a limitation of the present application.

[0044] In the description of the present application, “comprising” and “having” and any variants thereof cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, and can also include steps or units that are not listed but have similar features or functions or do not conflict with the listed steps or units or means.

[0045] Label misassignment, also known as index hopping or index misassignment or sample cross-talk, is a problem in the process of multiplex sequencing. In order to solve this problem to some extent or reduce the influence of this problem on sample detection to some extent, the embodiments of the present application propose a multiplex sample analysis method, please refer to Figure 1 The method comprises: S101 obtaining sequencing data of a multiplex sample, the sequencing data comprising reads from a plurality of samples; S102 splitting the sequencing data to obtain reads assigned to a specified sample; S103 filtering out corresponding part of reads belonging to a specified category of reads according to the classification result of the reads assigned to the specified sample by a preset classifier, comprising obtaining the values of the features of the reads assigned to the specified sample and determining the classification result of the reads assigned to the specified sample according to the values of the features by querying the preset classifier, wherein the preset classifier is a quantization scheme associated with the features and the probability of the reads belonging to the specified category of reads. The features referred to are a plurality of quantifiable features generated in the sequencing run and related to the signals for obtaining the reads assigned to the specified sample and / or assigning the reads to the specified sample, and then S104 detecting the specified sample according to the remaining reads assigned to the specified sample. The method can reduce the adverse effects of label hopping on sample detection.

[0046] Specifically, in step S101, sequencing data of a multiplex sample is obtained, and the sequencing data comprises reads from a plurality of samples. The sequencing data of the multiplex sample refers to sequencing data obtained by mixing the libraries of a plurality of samples in one reaction system for sequencing.

[0047] In order to sequence a plurality of samples in one reaction, the samples from different sources need to be distinguished in advance, i.e. different tags (index or barcode) are added to samples from different sources, which needs to be completed in the process of library construction. For example, taking the library construction of the second-generation sequencing platform of Illumina Company as an example, the main process includes: breaking the sample, e.g. DNA genome, into DNA fragments according to the Illumina / Solex DNA sample preparation method, then repairing the sticky ends formed by breaking into blunt ends; then adding a base“A” to the 3' end so that the DNA fragments can be connected to the adaptor with a“T” base at the 3' end and containing a tag sequence for marking the source of the sample; then using PCR technology to amplify the DNA fragments with adaptors at both ends and purifying the final PCR product to obtain a library containing a“tag” part and an“insert fragment” part. The specific process of library construction can also refer to the literature Multiplexed Illumina sequencing libraries from picogram quantities of DNA, Bowman et al. BMC Genomics 2013, 14: 466, http: / / www.biomedcentral.com / 1471-2164 / 14 / 466 and / or the literature Multiplex Illumina sequencing using DNA barcoding, Koon Ho Wong, Curr. Protoc. Mol. Biol. 101:7.11.1-7.11.11., 2013, https: / / currentprotocols.onlinelibrary.wiley.com / doi / 10.1002 / 0471142727.mb0711s101, the contents of which are all incorporated herein. Among them, the multiplex library is a mixture of sequencing libraries of a plurality of samples, and each sequencing library of the sample has a tag corresponding specifically to the sample. In some examples, the tag is a known sequence of greater than or equal to 1 bp, and the tags corresponding specifically to any two samples have a difference of at least one base. Further, in some examples, the tag is a known sequence of greater than or equal to 4 bp and less than 20 bp, and the tags corresponding specifically to any two samples have a difference of at least three bases to improve the fault tolerance.

[0048] When sequencing, the prepared libraries of multiple samples are mixed in one reaction system and sequenced simultaneously. For example, in the edge-sequencing based on surface fluorescence microscopic imaging, the reaction system generally includes sequencing primers, polymerase, and reaction substrates. Under the catalysis of DNA polymerase and under the conditions suitable for polymerization reaction, the (multiple) library is combined with the sequencing primer (the library or the amplified library or the library-primer complex is also called the nucleic acid molecule to be tested or the template), and based on the principle of base pairing, the reaction substrates such as nucleotides or their analogs are controllably connected to the end of the sequencing primer or the template, and the base at the corresponding position of the library / template is paired and connected, realizing the extension of the sequencing primer. The nucleotides or their analogs connected to the end of the sequencing primer are also called extension bases. The extension bases can be provided with optically detectable labels such as fluorescent groups, and different types of bases incorporated into the template (extension bases) can be associated or distinguished by different fluorescent groups or different light-emitting wave bands of the fluorescent groups, so that the type of the extension base can be determined by collecting the light-emitting signals generated by these labels during the extension of the sequencing primer by the microscopic imaging system, so as to obtain sequencing data. The sequencing of multiple samples can also refer to the literature Multiplex Illumina sequencing using DNA barcoding, Koon Ho Wong, Curr. Protoc. Mol. Biol. 101:7.11.1-7.11.11., 2013 https: / / currentprotocols.onlinelibrary.wiley.com / doi / 10.1002 / 0471142727.mb0711s101 and the literature https: / / www.illumina.com / documents / products / datasheets / datasheet_sequencing_multiplex.pdf and the literature Nature. 2008 Nov 6; 456(7218): 53-59, https: / / www.ncbi.nlm.nih.gov / pmc / articles / PMC 2581791 / , the contents of which are all incorporated herein.

[0049] In some examples, at the time of sequencing, a template can be constructed from one or a few rounds of imaging, where the template as referred to herein corresponds to a template in a biochemical reaction, a set of locations of chemical features (templates) on the surface, e.g., molecular clusters (amplified library) in an image; based on the images formed from the light signals collected by the microimaging system during the sequencing primer extension process of each round of reaction and the constructed template, the type of the extended base incorporated into the nucleic acid molecule under test in the round of reaction is identified to obtain sequencing data. The template construction and base identification process can be referred to PCT application WO2020037574A1, Chinese patent CN112823352B, and the literature Base-calling for next-generation sequencing platforms, Brief Bioinform, 2011 Sep; 12(5): 489-497, the contents of which are incorporated herein in their entirety.

[0050] In some examples, the multiplex library comprises the insert and the tag linked thereto, and the multiplex library is sequenced based on the surface fluorescent microimaging detection for sequencing-by-synthesis, comprising:

[0051] linking the multiplex library to the surface and amplifying the multiplex library on the surface, or amplifying the multiplex library and loading the amplified multiplex library to the surface, to form the surface with a plurality of amplicons;

[0052] subjecting the surface to conditions suitable for polymerization reaction, and performing multiple rounds of sequencing on the amplicons, comprising collecting images of each round of sequencing and detecting signals on the images corresponding to the locations of the amplicons, to identify the type of the base incorporated into the specified amplicon in the round of reaction, to determine the sequencing sequence of at least a portion of the insert and the sequencing sequence of the tag in the specified amplicon.

[0053] In certain embodiments, the amplification can be performed on the multiplex library in a solid phase or a liquid phase to obtain the amplified product of the multiplex library. Without limitation to the amplification method, for example, PCR (variable temperature amplification) can be performed using Taq enzyme or the like, RPA (recombinase polymerase amplification) can be performed using Bst or Bsu or recombinase or a multi-enzyme system or the like, isothermal amplification techniques such as RCA (rolling circle amplification), SDA (strand displacement amplification) and the like can be performed. For another example, bridge PCR or template walking amplification can be used to amplify the multiplex library on a solid phase surface to form amplicons on the surface; or RCA (rolling circle amplification) can be used to obtain the amplified product of the multiplex library in a liquid phase, and then the amplified product is carried to the surface to form DNA nanoballs or the like on the surface. (The multiplex library or the amplified multiplex library, without involving the need to distinguish the library before mixing, without affecting the understanding of those skilled in the art, is sometimes referred to as the library or template or nucleic acid molecule to be detected or amplified product or cluster or nanoball in this paper.)

[0054] In certain examples, when sequencing, the "tag" part and the "insert" part do not limit which part is read first and which part is read later, nor do they limit whether it is read once (for example, using one sequencing primer to read out the tag and the corresponding insert) or multiple times (for example, using two or more sequencing primers to read out one or more tags and the corresponding insert, respectively). For example, the library structure or one repeat unit of the library structure can be represented as 5' adapter-insert-tag-adapter, and at least one part of the 3' end adapter can be used to sequentially and continuously read out at least one part of the tag and the insert using a sequencing primer, and the tag part and the insert part in the read sequence can be distinguished according to the predetermined tag length and the tag being read before the insert; for another example, the library structure or one repeat unit of the library structure can be represented as 5' adapter-tag2-insert-tag1-adapter, and at least one part of the 3' end adapter can be used to sequentially and continuously read out at least one part of the tag1 and the insert using a sequencing primer 1, and at least one part of the tag2 and the insert can be read out using the same sequencing primer 2 as the 5' end adapter, and the tag (tag1 and tag2) part and the insert part can be determined based on the two read sequences; for another example, the library structure or one repeat unit of the library structure can be represented as 5' adapter-insert-X-tag-adapter, where X is a sequence-known spacer sequence, and the tag sequence can be determined using a sequencing primer 1 complementary to at least one part of the 3' end adapter, and a part of the insert can be determined using a sequencing primer 2 complementary to X. The specific process of sequencing can also refer to Chinese invention patent application CN113293205A, the contents of which are incorporated herein in their entirety.

[0055] The reads or readouts or readouts referred to herein, the sequencing sequences or sequencing sequences, a read is a sequencing readout of a polynucleotide sequence, sometimes refers to a readout containing index and insert (or part of the insert), sometimes refers to the readout of the insert (or part of the insert), sometimes refers to the readout of the index; in specific examples, those skilled in the art can clearly understand according to the relevant annotations and / or context. Generally, in multiplex sample detection analysis, if the sequencing data is split and assigned to each sample before, the reads generally contain tag readouts (tag readouts and corresponding insert readouts can be in a read, or in multiple reads that can determine the relative position relationship; the "corresponding insert" and "relative position relationship" referred to here means that the tag and the insert come from the same template), if the sequencing data is split and assigned to the specified sample for specified sample detection, then the reads generally refer to the insert readout without tag readout.

[0056] After obtaining the sequencing data of multiple samples, the sequencing data needs to be split according to step S102 to obtain the reads assigned to the specified sample. The sequencing data includes reads of multiple samples, so the sequencing data can be split to obtain reads assigned to the specified sample, which refers to the sample with a specified tag (barcode).

[0057] Specifically, the library of each sample contains connected inserts and tags corresponding specifically to the sample, the sequencing data includes reads of multiple samples, and the reads are split according to the part of the tag in the reads to obtain the sequencing data of each sample. Thus, the reads are split according to the specific tag sequence of the sample corresponding to the readout, and the reads assigned to the specified sample are obtained.

[0058] After obtaining the reads assigned to the specified sample, step S103 is performed, and according to the classification result of the pre-set classifier for the reads assigned to the specified sample, the corresponding part of the reads belonging to the specified category is filtered out.

[0059] The process includes obtaining the value of the feature of the reads assigned to the specified sample, and querying the pre-set classifier according to the value of the feature to determine the classification result of the reads assigned to the specified sample. The pre-set classifier is a quantitative scheme that associates the feature with the probability of the reads belonging to the specified category. The feature referred to is a plurality of quantifiable features generated in the sequencing run, which are related to the signals obtained when the reads assigned to the specified sample and / or the reads are assigned to the specified sample.

[0060] The values of the features of the split assigned reads of the designated sample are input into a preset classifier, which can determine the probability of the corresponding reads being a misassigned read (reads assigned to the designated sample by mistake) or a correctly assigned read (reads correctly assigned to the designated sample), which can be used to represent the confidence of the split being correct, and the higher the probability, the less likely the reads are assigned to the designated sample due to label jumping. Therefore, the reads assigned to the designated sample can be screened based on the probability score, for example, filtering out those reads determined by the preset classifier to have a higher probability of being a misassigned read, i.e. reads that the preset classifier considers to have a higher probability of being assigned by mistake due to label jumping.

[0061] In short, the preset classifier is a quantitative scheme for associating the features of the reads with the probability of the classification of the reads. In this article, the preset classifier is also sometimes referred to as a model or a function or a relationship. The model predicts the probability of the reads being misassigned reads by inputting the feature values of the reads, and optionally determines the category of the reads based on the probability; based on the predicted classification result, the reads are further screened or filtered, which can more accurately obtain the reads originating from the designated sample, and further more accurately detect and analyze the designated sample.

[0062] The features referred to are a plurality of quantifiable features generated in the sequencing run, which are related to the signals obtained for the assignment of the reads to the designated sample and / or the assignment of the reads to the designated sample.

[0063] The preset classifier can be established during the process of multiple sample analysis, such as in the process of S101 or S102 or before S103 based on training data, or it can be pre-trained to call the prediction result of the reads by the preset classifier for S103.

[0064] In some examples, the preset classifier is pre-trained based on a training data set for learning and training, and is stored for calling in multiple sample analysis.

[0065] In some examples, determining or constructing the preset classifier includes the following steps:

[0066] The training data set is obtained, which is sequencing data of multiple samples whose read assignment results and true value assignment results are known, and the training data set includes a plurality of reads containing at least a part of the sequence to be sequenced of each sample and the specific label sequence carried by the sample, and the training data set contains at least a certain number of reads;

[0067] The reads are divided into two categories according to whether the assignment result and the true value assignment result are consistent, and each read in the two categories is quantitatively represented in an alternative form based on the features of the reads.

[0068] The number of each type of reads contained in each alternative form is counted to establish the correspondence between the alternative form and the read type.

[0069] The training data set is sequencing data of multiple samples in which the sample assignment results and the true value assignment results of the reads are known, i.e., the training data set is sequencing data of multiple samples in which the normal reads (reads correctly assigned to the true source sample, sometimes also referred to as non-string reads in this paper) and abnormal reads (reads incorrectly assigned to a sample other than the true source sample, sometimes also referred to as mismatched reads or hopping-reads or jumping reads or string reads in this paper) are known. The more data contained in the training data set, the more accurate the prediction of the preset classifier for the type of reads. Therefore, in the embodiments of the present application, the training data set contains at least a certain number of reads, such as 10,000 reads, 100,000 reads or even more reads. The reads can be divided into two categories according to whether the assignment result and the true value assignment result are consistent, i.e., into normal reads and abnormal reads. Then each read in the two types of reads is quantitatively represented as an alternative form according to the characteristics of the normal reads and the abnormal reads, thereby establishing the correspondence between the alternative form and the read type, so that the type of the read can be determined according to the alternative form when the prediction classifier is applied subsequently. Further, in order to accurately determine the type of the read according to the alternative form when the prediction classifier is applied subsequently, the number of each type of reads contained in each alternative form can also be counted to establish the correspondence between the alternative form and the read type.

[0070] In some examples, the training data set and the unsplit sequencing data in the multiple sample analysis method in any example can be obtained by using the same or different sequencing methods, sequencing platforms and / or sequencing modes. The sequencing methods include but are not limited to sequencing by synthesis, sequencing by ligation or pyrosequencing; the sequencing platforms include but are not limited to the Illunima / solexa platform, the ABI Solid platform and the Roche 454 platform; the sequencing modes can use common sequencing modes of commercially available sequencing platforms such as SE-double index (single-end double index), SE-single index (single-end single index), PE-single index (double-end single index), PE-double index (double-end double index) and the like. Preferably, the training data set and the unsplit sequencing data in the multiple sample analysis method in any example come from the same sequencing method or the same manufacturer's sequencing platform, which facilitates more accurate multiple sample analysis. Further, sequencing the multiple samples on different sequencers of the same sequencing platform can ensure that the training data set obtained does not have too much bias.

[0071] The inventors generate features of tens of reads based on the information of images in the sequencing process, the identified bases, the sequence characteristics of the reads, etc. After screening and testing, some features are determined which can quantitatively represent the reads and have obvious differences in the distribution in the two types of reads, and are used to construct the preset classifier.

[0072] In some examples, the features of the reads are selected from at least one of the following a)-b): a) the signal intensity of the insert portion readout or any feature having a Spearman correlation coefficient with the signal intensity of the insert portion readout not less than 0.5; b) the signal intensity of the tag readout or any feature having a Spearman correlation coefficient with the signal intensity of the tag readout not less than 0.5.

[0073] In some examples, the features of the reads include a) and b); more specifically, the features of the reads include the signal intensity of the insert portion readout (hereinafter sometimes represented as readInts) and the signal intensity of the tag readout (hereinafter sometimes represented as barInts). The value of the feature of the signal intensity of the insert portion readout readInts can be represented by the average or median of the signal intensity of all or part of the bases of the insert portion readout on the corresponding image. The value of the feature of the signal intensity of the tag readout barInts can be represented by the average or median of the signal intensity of all the bases of the tag readout on the corresponding image.

[0074] The determination of the position of the extended or readout base on the corresponding image can be determined by using conventional methods such as the center of gravity method, etc. The signal intensity of the extended or readout base on the image can be absolute signal intensity or relative signal intensity. If not specifically stated, it generally refers to relative signal intensity. For example, the signal intensity after image preprocessing such as background removal (if it is a grayscale image, it can be equivalent to the grayscale value or pixel at the specified position, and if it is a color image, it can be converted into a grayscale image for calculation in order to quickly determine), such as the signal intensity after phasing, pre-phasing and / or crosstalk correction, etc. The signal intensity at the corresponding position can be determined by using the interpolation method. In the sequencing process based on the detection and identification of the extended base by fluorescence signal, the signal intensity is also sometimes referred to as brightness. In the sequencing platform based on surface fluorescence microscopic imaging detection, the signal intensity of a position corresponding to a chemical feature on the image collected (such as the bright spot feature presented on the image) in each round of reaction for the edge synthesis sequencing can be represented as an array of four bases, for example, represented as [B1, B2, B3, B4], wherein B1 to B4 are respectively the signals generated on A, T, C and G. For example, the signal intensity of a certain chemical feature position obtained in a round of reaction is [80, 10, 5, 5], and it is determined that the extended base type corresponding to the chemical feature position in this round of reaction is A, wherein 80 is the maximum signal intensity (or maximum brightness value; or maximum grayscale value; or maximum Ints), and 10 is the second largest signal intensity (or second largest brightness value; or second largest grayscale value; or second largest Ints). In a certain example, the insert portion readout is not less than 50 bp, and the maximum signal intensity of 16 to 25 bases on the corresponding image from the relative middle, for example, the 15th base, is sorted, and the median value represents the value of the readInts; the tag readout is not less than 6 bp, and the maximum signal intensity of each base on the corresponding image of the tag readout is sorted, and the median value represents the value of the barInts. In this way, the whole readout situation can be better reflected, and the value of the feature can also be quickly obtained. In some examples, the features of the read segment in addition to a) and b) also include the signal intensity of the specific base of the tag readout or any feature with a Spearman correlation coefficient with the signal intensity of the specific base of the tag readout not less than 0.5.In certain examples, the signal intensity of a particular base of a tag readout includes the maximum signal intensity of the particular base of the tag readout and / or the second largest signal intensity of the particular base of the tag readout, i.e., in addition to a), b), the features of the read segment also include one or more of c) and d): c) the maximum signal intensity of the particular base of the tag readout or any feature having a Spearman correlation coefficient with the maximum signal intensity of the particular base of the tag readout that is not less than 0.5; d) the second largest signal intensity of the particular base of the tag readout or any feature having a Spearman correlation coefficient with the second largest signal intensity of the particular base of the tag readout that is not less than 0.5. The maximum signal intensity of the particular base is sometimes denoted as badQMInts; the second largest signal intensity of the particular base is sometimes denoted as badQSInts. The particular base is, for example, the first base of the inverse Q value, the second base of the inverse Q value, or the third base of the inverse Q value.

[0075] In certain examples, the maximum signal intensity of a particular base of a tag readout or any feature having a Spearman correlation coefficient with the maximum signal intensity of the particular base of the tag readout that is not less than 0.5; and the second largest signal intensity of the particular base of the tag readout or any feature having a Spearman correlation coefficient with the second largest signal intensity of the particular base of the tag readout that is not less than 0.5 include:

[0076] the maximum signal intensity of the first base of the inverse Q value of a tag readout or any feature having a Spearman correlation coefficient with the maximum signal intensity of the first base of the inverse Q value of the tag readout that is not less than 0.5; and,

[0077] the second largest signal intensity of the first base of the inverse Q value of a tag readout or any feature having a Spearman correlation coefficient with the second largest signal intensity of the first base of the inverse Q value of the tag readout that is not less than 0.5.

[0078] In certain examples, the maximum signal intensity of a particular base of a tag readout or any feature having a Spearman correlation coefficient with the maximum signal intensity of the particular base of the tag readout that is not less than 0.5; and the second largest signal intensity of the particular base of the tag readout or any feature having a Spearman correlation coefficient with the second largest signal intensity of the particular base of the tag readout that is not less than 0.5 include:

[0079] the maximum signal intensity of the second base of the inverse Q value of a tag readout or any feature having a Spearman correlation coefficient with the maximum signal intensity of the second base of the inverse Q value of the tag readout that is not less than 0.5; and,

[0080] the second largest signal intensity of the second base of the inverse Q value of a tag readout or any feature having a Spearman correlation coefficient with the second largest signal intensity of the second base of the inverse Q value of the tag readout that is not less than 0.5.

[0081] In some examples, the maximum signal intensity of a specific base read by a tag or any feature with a Spearman correlation coefficient with the maximum signal intensity of the specific base read by the tag not less than 0.5; and the second maximum signal intensity of the specific base read by the tag or any feature with a Spearman correlation coefficient with the second maximum signal intensity of the specific base read by the tag not less than 0.5 include:

[0082] the maximum signal intensity of the third base from the bottom in terms of Q value read by a tag or any feature with a Spearman correlation coefficient with the maximum signal intensity of the third base from the bottom in terms of Q value read by the tag not less than 0.5; and

[0083] the second maximum signal intensity of the third base from the bottom in terms of Q value read by a tag or any feature with a Spearman correlation coefficient with the second maximum signal intensity of the third base from the bottom in terms of Q value read by the tag not less than 0.5.

[0084] In some examples, a quality score is assigned to each read base on a tag, i.e., a base quality value (Q value) is calculated for each read base to determine the reliability of each read base. The Q values of the read bases are sorted in descending order, and the maximum signal intensity and the second maximum signal intensity of the first, second or third base from the bottom in terms of Q value are taken as the feature values of badQMInts and badQSInts of the tag read, respectively. In this way, the overall situation of the read can be better reflected, and the feature values can be quickly obtained.

[0085] In some examples, the maximum signal intensity and the second maximum signal intensity of the second base from the bottom in terms of Q value are more stable than the maximum signal intensity and the second maximum signal intensity of the first base from the bottom in terms of Q value or the maximum signal intensity and the second maximum signal intensity of the third base from the bottom in terms of Q value, which can enhance the universality and practicality of the trained classifier.

[0086] In some examples, the feature values readInts, barInts, badQMInts and badQSInts are obtained by extracting the gray value or brightness value of the bright spot on the image based on surface fluorescence microscopic imaging detection during sequencing, and can be selected based on actual application requirements among the above predetermined features. Generally, a combination of the two feature values readInts and barInts can be selected. In some examples, in the case of a short tag sequence and high sequence similarity between tags, a combination of the four feature values readInts, barInts, badQMInts and badQSInts can be selected.

[0087] In some embodiments, quantifying each read in the two types of reads based on the features of the reads into an alternative form means a form identified by the features corresponding to the features of the processed reads, such as a normalized feature value, or a feature value revalued based on a weight coefficient corresponding to the feature value, and the like. Taking the normalized representation as an example, quantifying each read in the two types of reads based on the features of the reads into an alternative form includes: performing normalization processing on the feature values corresponding to the features of the reads to obtain normalized feature values; and quantifying the normalized feature values into an alternative form.

[0088] In some examples, quantifying each read in the two types of reads into an alternative form according to the features of the normal reads and the abnormal reads includes:

[0089] performing normalization processing on the feature values corresponding to the features of the normal reads and the abnormal reads to obtain normalized feature values, and quantifying the normalized feature values into an alternative form.

[0090] Specifically, performing normalization processing on the four feature values readInts, barInts, badQMInts and badQSInts of each normal read and abnormal read to form four-dimensional discrete data from [0, 0, 0, 0] to [9, 9, 9, 9], and representing the discrete data formed after the normalization processing into an alternative form; or,

[0091] performing normalization processing on the two feature values readInts and barInts of each normal read and abnormal read to form two-dimensional discrete data from [0, 0] to [9, 9], and representing the discrete data formed after the normalization processing into an alternative form.

[0092] In some examples, the preset classifier constructed is shown in Table 1:

[0093] Table 1

[0094]

[0095]

[0096] wherein Type represents the discrete data formed after normalization of the feature values, Type is represented by Ai, i is the number of rows of the current table, for example, Ai can be 2239.

[0097] HoppedNum represents the number of abnormal reads in the sequencing data, for example, in Table 1, it is represented by Bi, i is the number of rows of the current table, for example, Bi can be 0, i.e. the number of abnormal reads is 0.

[0098] TotalNum represents the total number of reads in the sequencing data, denoted as Ci in Table 1, i is the row number of the current table, for example, Ci can be 29.

[0099] Score represents the probability, the value of -log(hoppedNum / totalNum) is rounded, the value range is [0, 9], which is used to represent the confidence of correct splitting, the higher the probability, the smaller the possibility of read being incorrectly allocated to the specified sample due to label jumping; on the contrary, the lower the probability, the greater the possibility of read being incorrectly allocated to the specified sample due to label jumping.

[0100] In actual application process, after multiple sample data are counted, a classification table of label jumping reads is generated, and a preset classifier can be formed based on the data in the classification table.

[0101] After the read segments that help the specified category in the read segments allocated to the specified sample are filtered out according to the preset classifier, the specified sample is detected according to the remaining read segments allocated to the specified sample according to step S104.

[0102] The remaining read segments are those read segments left after the read segments predicted by the preset classifier to be incorrectly allocated to the specified sample with high probability are removed, and then the specified sample is detected according to the remaining read segments, so as to obtain the biological information of the specified sample, for example, the read segments after screening are used to detect the mutation of the biological body from which the specified sample is derived. In this way, the influence of read segment error splitting caused by label jumping can be reduced to a certain extent, and the accuracy of detecting the specified sample based on the read segments allocated to the specified sample is improved. Specifically, the biological information of the specified sample can be obtained by detecting the specified sample according to the remaining read segments allocated to the specified sample, for example, the base sequence order and base mutation information can be obtained.

[0103] In the embodiments of the present application, a multiple sample analysis device is also provided, as shown in Figure 2 , comprising:

[0104] The acquisition unit 201 is configured to acquire sequencing data of multiple samples, wherein the sequencing data comprises read segments from multiple samples.

[0105] The splitting unit 202 is configured to split the sequencing data to obtain read segments allocated to the specified sample.

[0106] The filtering unit 203 is configured to filter out a corresponding part of reads belonging to a specified category from the reads assigned to a specified sample according to a classification result of the reads assigned to the specified sample by a preset classifier, including obtaining a value of a feature of the reads assigned to the specified sample, and querying the preset classifier according to the value of the feature to determine the classification result of the reads assigned to the specified sample,

[0107] The preset classifier is a quantization scheme in which the feature is associated with a probability of the reads belonging to the specified category.

[0108] The detection unit 204 is configured to detect the specified sample according to the remaining reads assigned to the specified sample.

[0109] In some examples, the obtaining unit includes:

[0110] The constructing sub-unit is configured to construct a multiplex library containing a plurality of samples, the multiplex library being a mixture of sequencing libraries of the plurality of samples, each sequencing library of each sample being labeled with a label corresponding to the sample in a specific manner, the label being a known sequence of greater than or equal to 1 bp, and the labels corresponding to any two samples in a specific manner being different by at least one base; and the sequencing sub-unit is configured to sequence the multiplex library.

[0111] In some examples, the sequencing data further includes a sequencing sequence of the label corresponding to the specified sample in a specific manner, and the sequencing data is split according to the sequencing sequence of the label to obtain reads assigned to the specified sample.

[0112] In some examples, the multiplex library contains the insert fragments and the labels connected thereto, and the obtaining unit is specifically configured to sequence the multiplex library by sequencing-by-synthesis based on surface fluorescence microscopic imaging detection, and the obtaining unit includes:

[0113] The amplifying unit is configured to connect the multiplex library to a surface and amplify the multiplex library on the surface, or amplify the multiplex library and load the amplified multiplex library to the surface, to form a surface with a plurality of amplicons;

[0114] The multiple-round sequencing sub-unit is configured to place the surface in a condition suitable for polymerization reaction, sequence the amplicons in multiple rounds, including collecting images of each round of sequencing and detecting signals of positions of the corresponding amplicons on the images, to identify types of bases incorporated into a specified amplicon in the round of reaction, to determine a sequencing sequence of at least part of the insert fragments and a sequencing sequence of the label in the specified amplicon.

[0115] In some examples, the feature is selected from at least one of a) to d) below:

[0116] a) a signal intensity of a sequencing sequence of the insert or any feature having a Spearman correlation coefficient with the signal intensity of the sequencing sequence of the insert not less than 0.5;

[0117] b) a signal intensity of a sequencing sequence of the tag or any feature having a Spearman correlation coefficient with the signal intensity of the sequencing sequence of the tag not less than 0.5;

[0118] c) a maximum signal intensity of a specific base of a tag read or any feature having a Spearman correlation coefficient with the maximum signal intensity of the specific base of the tag read not less than 0.5;

[0119] d) a second maximum signal intensity of a specific base of a tag read or any feature having a Spearman correlation coefficient with the second maximum signal intensity of the specific base of the tag read not less than 0.5.

[0120] In some examples, the apparatus further comprises a classifier determining unit, which comprises:

[0121] a data obtaining sub-unit configured to obtain a training data set, the training data set being sequencing data of multiple samples whose reads and the samples' allocation results are known, and the training data set comprising a plurality of reads comprising at least a part of a sequence to be sequenced of each sample and a specific tag sequence carried by the sample, and the training data set comprising at least a specific number of reads;

[0122] a representation sub-unit configured to divide the reads into two categories according to whether the allocation results and the true value allocation results are consistent, and to represent each read in the two categories in an alternative form based on a feature of the read;

[0123] a establishing sub-unit configured to count the number of each category of reads contained in each alternative form, to establish a correspondence between the alternative form and the category of reads.

[0124] In some examples, the feature of the read comprises at least one of:

[0125] a signal intensity of a part of an insert read or any feature having a Spearman correlation coefficient with the signal intensity of the insert read not less than 0.5;

[0126] a signal intensity of a tag read or any feature having a Spearman correlation coefficient with the signal intensity of the tag read not less than 0.5;

[0127] The maximum signal intensity of a specific base read by the tag or any feature whose Spearman correlation coefficient with the maximum signal intensity of a specific base read by the tag is not less than 0.5;

[0128] The second largest signal intensity of a specific base read by a tag or any feature whose Spearman correlation coefficient with the second largest signal intensity of a specific base read by a tag is not less than 0.5.

[0129] In some examples, the maximum signal intensity of a specific base read by the tag or any feature having a Spearman correlation coefficient with the maximum signal intensity of a specific base read by the tag is not less than 0.5; and

[0130] The second largest signal intensity of a specific base read by the tag, or any feature whose Spearman correlation coefficient with the second largest signal intensity of a specific base read by the tag is not less than 0.5, including:

[0131] The maximum signal intensity of the base with the lowest Q value read from the tag, or any feature whose Spearman correlation coefficient with the maximum signal intensity of the base with the lowest Q value read from the tag is not less than 0.5; and

[0132] The second largest signal intensity of the base with the lowest Q value read from the tag, or any feature whose Spearman correlation coefficient with the second largest signal intensity of the base with the lowest Q value read from the tag is not less than 0.5.

[0133] In some examples, the maximum signal intensity of a specific base read by the tag or any feature having a Spearman correlation coefficient with the maximum signal intensity of a specific base read by the tag is not less than 0.5; and

[0134] The second largest signal intensity of a specific base read by the tag, or any feature whose Spearman correlation coefficient with the second largest signal intensity of a specific base read by the tag is not less than 0.5, including:

[0135] The maximum signal intensity of the penultimate base with the Q value read from the tag, or any feature having a Spearman correlation coefficient of not less than 0.5 with the maximum signal intensity of the penultimate base with the Q value read from the tag; and

[0136] The second largest signal intensity of the penultimate base Q value read from the tag or any feature whose Spearman correlation coefficient with the second largest signal intensity of the penultimate base Q value read from the tag is not less than 0.5.

[0137] In some examples, the maximum signal intensity of a specific base read by the tag or any feature having a Spearman correlation coefficient with the maximum signal intensity of a specific base read by the tag is not less than 0.5; and

[0138] Any feature of the second largest signal intensity of a specific base read by the tag or a Spearman correlation coefficient of the second largest signal intensity of a specific base read by the tag not less than 0.5, including:

[0139] Any feature of the third largest signal intensity of a specific base read by the tag or a Spearman correlation coefficient of the third largest signal intensity of a specific base read by the tag not less than 0.5, including:

[0140] Any feature of the second largest signal intensity of a specific base read by the tag or a Spearman correlation coefficient of the second largest signal intensity of a specific base read by the tag not less than 0.5, including:

[0141] In some examples, the representation subunit is configured to:

[0142] The feature value corresponding to the feature of the read is normalized to obtain a normalized feature value;

[0143] The normalized feature value is quantitatively represented in an alternative form.

[0144] It should be noted that the specific implementation of each unit and subunit in this embodiment can refer to the corresponding content in the foregoing, which will not be described in detail here.

[0145] In another embodiment of the present application, a computing device is also provided, which comprises:

[0146] A memory for storing a computer executable program;

[0147] One or more processors for executing the computer executable program and implementing the multiple sample analysis method according to any one of the above.

[0148] In another embodiment of the present application, a computer readable storage medium for storing a computer executable program is also provided, and executing the program includes completing the multiple sample analysis method according to any one of the above.

[0149] It should be noted that the specific implementation of the processor in this embodiment can refer to the corresponding content in the foregoing, which will not be described in detail here.

[0150] In still another embodiment of the present application, a method for constructing a classifier capable of predicting whether a read of mixed sample sequencing data split to a specified sample is a string sample read or a non-string sample read is also provided, and the method for constructing the classifier comprises:

[0151] obtaining a training data set, the training data set being sequencing data of multiple samples of reads of which sample assignment results and true value assignment results are known, the training data set comprising a plurality of reads comprising at least a part of a sequence to be sequenced of each sample and a specific tag sequence carried by the sample, the training data set comprising at least a specific number of reads;

[0152] classifying the reads into two categories according to whether the assignment results and the true value assignment results are consistent, and quantitatively representing each read in the two categories of reads as an alternative form based on the characteristics of the reads;

[0153] counting the number of each category of reads contained in each alternative form to establish a correspondence between the alternative forms and the categories of reads;

[0154] based on the correspondence, constructing a classifier.

[0155] The training data set is known in which the reads and the samples from which the reads originate, i.e., the training data set is a sample of which the label can be known, even if a label jump is generated, it can be identified. The more the number of training data sets, the higher the accuracy of the analysis based on the statistical results, therefore, in the embodiments of the present application, the training data set comprises at least a specific number of reads, for example, 10,000 reads, 100,000 reads or even more reads. The true value assignment result represents the reads that generate a label jump and the reads that do not generate a label jump. The assignment results and the true value assignment results are consistent, i.e., the reads are classified into two categories, i.e., the reads that generate a label jump and the reads that do not generate a label jump. Then, according to the characteristics of the reads, each read in the two categories of reads is quantitatively represented as an alternative form, so that the type of the read can be determined according to the alternative form. Further, in order to enable the alternative form to be accurately applied to subsequent read category identification, the number of each category of reads contained in each alternative form can be counted to establish a correspondence between the alternative forms and the categories of reads.

[0156] In some examples, the characteristics of the reads include at least one of the following:

[0157] any feature of which the signal intensity of the insert portion readout or the signal intensity of the insert readout has a Spearman correlation coefficient of not less than 0.5;

[0158] any feature of which the signal intensity of the tag readout or the signal intensity of the tag readout has a Spearman correlation coefficient of not less than 0.5;

[0159] any feature of which the maximum signal intensity of the specific base of the tag readout or the maximum signal intensity of the specific base of the tag readout has a Spearman correlation coefficient of not less than 0.5;

[0160] The second largest signal intensity of a specific base read by the tag or any feature with a Spearman correlation coefficient with the second largest signal intensity of the specific base read by the tag not less than 0.5.

[0161] In certain examples, the largest signal intensity of a specific base read by the tag or any feature with a Spearman correlation coefficient with the largest signal intensity of the specific base read by the tag not less than 0.5; and,

[0162] The second largest signal intensity of a specific base read by the tag or any feature with a Spearman correlation coefficient with the second largest signal intensity of the specific base read by the tag not less than 0.5, including:

[0163] The largest signal intensity of the first base with the smallest reciprocal Q value read by the tag or any feature with a Spearman correlation coefficient with the largest signal intensity of the first base with the smallest reciprocal Q value read by the tag not less than 0.5; and,

[0164] The second largest signal intensity of the first base with the smallest reciprocal Q value read by the tag or any feature with a Spearman correlation coefficient with the second largest signal intensity of the first base with the smallest reciprocal Q value read by the tag not less than 0.5.

[0165] In certain examples, the largest signal intensity of a specific base read by the tag or any feature with a Spearman correlation coefficient with the largest signal intensity of the specific base read by the tag not less than 0.5; and,

[0166] The second largest signal intensity of a specific base read by the tag or any feature with a Spearman correlation coefficient with the second largest signal intensity of the specific base read by the tag not less than 0.5, including:

[0167] The largest signal intensity of the second base with the smallest reciprocal Q value read by the tag or any feature with a Spearman correlation coefficient with the largest signal intensity of the second base with the smallest reciprocal Q value read by the tag not less than 0.5; and,

[0168] The second largest signal intensity of the second base with the smallest reciprocal Q value read by the tag or any feature with a Spearman correlation coefficient with the second largest signal intensity of the second base with the smallest reciprocal Q value read by the tag not less than 0.5.

[0169] In certain examples, the largest signal intensity of a specific base read by the tag or any feature with a Spearman correlation coefficient with the largest signal intensity of the specific base read by the tag not less than 0.5; and,

[0170] The second largest signal intensity of a specific base read by the tag or any feature with a Spearman correlation coefficient with the second largest signal intensity of the specific base read by the tag not less than 0.5, including:

[0171] a maximum signal intensity of a base ranked third from the bottom in a Q value of tag reading or any feature with a Spearman correlation coefficient with the maximum signal intensity of the base ranked third from the bottom in the Q value of tag reading not less than 0.5; and

[0172] a second maximum signal intensity of a base ranked third from the bottom in a Q value of tag reading or any feature with a Spearman correlation coefficient with the second maximum signal intensity of the base ranked third from the bottom in the Q value of tag reading not less than 0.5.

[0173] In some examples, in order to make the data quantization more accurate, the alternative form can be determined based on normalization. In the embodiments of the present application, and based on the predetermined features, quantizing each read in the two types of reads into the alternative form includes: performing normalization on the feature values corresponding to the predetermined features to obtain normalized feature values; and quantizing the normalized feature values into the alternative form.

[0174] In another embodiment of the present application, a computer readable storage medium is also provided for storing a program for computer execution, and execution of the program includes completing the method of constructing the classifier as described above.

[0175] In still another embodiment of the present application, a computer device is also provided, which includes:

[0176] a memory for storing a computer executable program;

[0177] one or more processors for executing the computer executable program and implementing the method of constructing the classifier as described above.

[0178] In another embodiment of the present application, an application method of the classifier is also provided, which includes:

[0179] classifying the reads assigned to the specified sample to obtain a classification result;

[0180] so that according to the classification result, the corresponding part of reads belonging to the specified category is filtered out, including obtaining the values of features of the reads assigned to the specified sample, and querying the preset classifier according to the values of the features to determine the classification result of the reads assigned to the specified sample; and detecting the specified sample according to the remaining reads assigned to the specified sample;

[0181] wherein the reads assigned to the specified sample are obtained by splitting sequencing data of multiple samples, and the sequencing data includes reads from multiple samples.

[0182] The multiple sample analysis method or device disclosed in the embodiments of the present application obtains the classification result of the reads assigned to the specified sample through a preset classifier, and filters out the reads belonging to the specified category, for example, the reads that are more likely to cause label jumping and are wrongly assigned to the specified sample, thereby reducing the influence of label jumping and wrongly assigned reads on the detection of the specified sample.

[0183] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part.

[0184] The skilled person can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in general terms. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0185] The steps of the method or algorithm described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art.

[0186] The above description of the disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multiplex sample analysis method, characterized by, The method comprises: obtaining sequencing data of multiple samples, the sequencing data comprising reads from multiple samples; splitting the sequencing data to obtain reads assigned to a specified sample; filtering out corresponding part of reads belonging to a specified category from the reads assigned to the specified sample according to a classification result of the reads assigned to the specified sample by a preset classifier, comprising obtaining a value of a feature of the reads assigned to the specified sample, and querying the preset classifier according to the value of the feature to determine the classification result of the reads assigned to the specified sample, wherein the preset classifier is a quantization scheme associated with the feature and the probability of the reads assigned to the specified sample belonging to the specified category, and the feature is a plurality of quantifiable features generated in a sequencing run and related to signals for obtaining the reads assigned to the specified sample and / or assigning the reads to the specified sample; and detecting the specified sample according to the remaining reads assigned to the specified sample.

2. The method of claim 1, wherein, The method comprises: constructing a multiplex library comprising multiple samples, the multiplex library being a mixture of sequencing libraries of multiple samples, each sequencing library of a sample being labeled with a label specific to the sample, the label being a known sequence of greater than or equal to 1 bp, and the labels specific to any two samples being different by at least one base; and sequencing the multiplex library; optionally, the sequencing data further comprises a sequencing sequence of the label specific to the specified sample, and the sequencing data is split according to the sequencing sequence of the label to obtain reads assigned to the specified sample; optionally, the multiplex library comprises an insert and the label connected thereto, and the multiplex library is sequenced by sequencing-by-synthesis based on surface fluorescence microscopic imaging detection, comprising: connecting the multiplex library to a surface and amplifying the multiplex library on the surface, or amplifying the multiplex library and loading the amplified multiplex library to the surface, to form a surface with multiple amplicons; placing the surface under conditions suitable for polymerization reaction, and performing multiple rounds of sequencing on the amplicons, comprising collecting images of each round of sequencing and detecting signals of positions of the amplicons on the images to identify types of bases incorporated into a specified amplicon in the round of reaction, to determine a sequencing sequence of at least part of the insert and a sequencing sequence of the label in the specified amplicon; optionally, the feature is selected from at least one of the following a)-d): a) signal intensity of the sequencing sequence of the insert or any feature having a Spearman correlation coefficient with the signal intensity of the sequencing sequence of the insert not less than 0.5; b) signal intensity of the sequencing sequence of the label or any feature having a Spearman correlation coefficient with the signal intensity of the sequencing sequence of the label not less than 0.5; c) maximum signal intensity of a specific base read by the label or any feature having a Spearman correlation coefficient with the maximum signal intensity of the specific base read by the label not less than 0.5; d) second maximum signal intensity of a specific base read by the label or any feature having a Spearman correlation coefficient with the second maximum signal intensity of the specific base read by the label not less than 0.

5.

3. The method according to claim 1 or 2, characterized in that, The preset classifier is determined by the following steps: obtaining a training data set, the training data set being sequencing data of multiple samples of which the samples of reads are assigned results and true value assignment results, the training data set comprising a plurality of reads comprising at least a part of the sequence to be sequenced of each sample and a specific tag sequence carried by the sample, the training data set comprising at least a certain number of reads; According to whether the assignment result and the true value assignment result are consistent, the reads are divided into two categories, and each read in the two categories of reads is quantitatively represented in an alternative form based on the characteristics of the reads; The number of reads of each category contained in each alternative form is counted to establish the correspondence between the alternative form and the read category; Optionally, the characteristics of the reads include at least one of the following: The signal intensity of the insert portion read or the signal intensity of the insert read has a Spearman correlation coefficient not less than 0.5; The signal intensity of the tag read or any feature with a Spearman correlation coefficient not less than 0.5 with the signal intensity of the tag read; The maximum signal intensity of a specific base of the tag read or any feature with a Spearman correlation coefficient not less than 0.5 with the maximum signal intensity of the specific base of the tag read; The second largest signal intensity of a specific base of the tag read or any feature with a Spearman correlation coefficient not less than 0.5 with the second largest signal intensity of the specific base of the tag read; Optionally, the maximum signal intensity of a specific base of the tag read or any feature with a Spearman correlation coefficient not less than 0.5 with the maximum signal intensity of the specific base of the tag read; and, The second largest signal intensity of a specific base of the tag read or any feature with a Spearman correlation coefficient not less than 0.5 with the second largest signal intensity of the specific base of the tag read, includes: The maximum signal intensity of the first base with the smallest reciprocal of Q value of the tag read or any feature with a Spearman correlation coefficient not less than 0.5 with the maximum signal intensity of the first base with the smallest reciprocal of Q value of the tag read; and, the second largest signal intensity of the first base with the smallest reciprocal of Q value of the tag read or any feature with a Spearman correlation coefficient not less than 0.5 with the second largest signal intensity of the first base with the smallest reciprocal of Q value of the tag read; Optionally, the maximum signal intensity of a specific base of the tag read or any feature with a Spearman correlation coefficient not less than 0.5 with the maximum signal intensity of the specific base of the tag read; and, The second largest signal intensity of a specific base of the tag read or any feature with a Spearman correlation coefficient not less than 0.5 with the second largest signal intensity of the specific base of the tag read, includes: The maximum signal intensity of the second base with the smallest reciprocal of Q value of the tag read or any feature with a Spearman correlation coefficient not less than 0.5 with the maximum signal intensity of the second base with the smallest reciprocal of Q value of the tag read; and, The second largest signal intensity of the second base with the smallest reciprocal of Q value of the tag read or any feature with a Spearman correlation coefficient not less than 0.5 with the second largest signal intensity of the second base with the smallest reciprocal of Q value of the tag read; Optionally, the maximum signal intensity of the specific base of the tag readout or any feature with a Spearman correlation coefficient with the maximum signal intensity of the specific base of the tag readout not less than 0.5; and The second largest signal intensity of the specific base of the tag readout or any feature with a Spearman correlation coefficient with the second largest signal intensity of the specific base of the tag readout not less than 0.5, including: The maximum signal intensity of the third base from the bottom of the Q value of the tag readout or any feature with a Spearman correlation coefficient with the maximum signal intensity of the third base from the bottom of the Q value of the tag readout not less than 0.5; and The second largest signal intensity of the third base from the bottom of the Q value of the tag readout or any feature with a Spearman correlation coefficient with the second largest signal intensity of the third base from the bottom of the Q value of the tag readout not less than 0.

5.

4. The method of claim 3, wherein, The alternative form of the quantized representation of each read in the two categories of reads based on the read-based features includes: The feature value corresponding to the read-based feature is normalized to obtain a normalized feature value; The normalized feature value is quantized and represented in an alternative form.

5. A multiple sample analysis device, characterized by, It includes: An acquisition unit is configured to acquire sequencing data of a multiplex sample, wherein the sequencing data includes reads from a plurality of samples; A splitting unit is configured to split the sequencing data to obtain reads assigned to a specified sample; A filtering unit is configured to filter out a corresponding part of reads belonging to a specified category of reads according to a classification result of the reads assigned to the specified sample by a preset classifier, including obtaining a value of a feature of the reads assigned to the specified sample, and querying the preset classifier according to the value of the feature to determine the classification result of the reads assigned to the specified sample, Wherein, the preset classifier is a quantization scheme associated with the feature and the probability of the reads belonging to the specified category of reads, and the feature is a plurality of quantifiable features generated during sequencing and related to the signal for obtaining the reads assigned to the specified sample and / or assigning the reads to the specified sample; And a detection unit is configured to detect the specified sample according to the remaining reads assigned to the specified sample.

6. The apparatus of claim 5, wherein, The acquisition unit includes: A construction subunit is configured to construct a multiplex library containing a plurality of samples, wherein the multiplex library is a mixture of sequencing libraries of a plurality of samples, each sequencing library of a sample has a label corresponding specifically to the sample, the label is a known sequence of greater than or equal to 1 bp, and the labels corresponding specifically to any two samples have at least one base difference; and A sequencing subunit is configured to sequence the multiplex library; Optionally, the sequencing data further includes a sequencing sequence of a label corresponding specifically to a specified sample, and the sequencing data is split to obtain reads assigned to the specified sample according to the sequencing sequence of the label; Optionally, the multiplex library contains connected insert fragments and the labels, and the acquisition unit is specifically configured to sequence the multiplex library by sequencing-by-synthesis based on surface fluorescence microscopic imaging detection, and the acquisition unit includes: an amplification unit configured to ligate the multiplex library to a surface and amplify the multiplex library on the surface, or to amplify the multiplex library and load the amplified multiplex library to the surface, to form a surface with a plurality of amplicons; a multi-round sequencing subunit configured to subject the surface to a plurality of rounds of sequencing under conditions suitable for polymerization, to perform a plurality of rounds of sequencing on the amplicons, including acquiring images of each round of sequencing and detecting signals of locations of the amplicons on the images to identify a type of a base incorporated into a given amplicon in the round of sequencing, to determine a sequencing sequence of at least a portion of the insert in the given amplicon and a sequencing sequence of the tag; optionally, the feature is selected from at least one of a)-d): a) a signal intensity of the sequencing sequence of the insert or any feature with a Spearman correlation coefficient with the signal intensity of the sequencing sequence of the insert no less than 0.5; b) a signal intensity of the sequencing sequence of the tag or any feature with a Spearman correlation coefficient with the signal intensity of the sequencing sequence of the tag no less than 0.5; c) a maximum signal intensity of a specific base of the tag read or any feature with a Spearman correlation coefficient with the maximum signal intensity of the specific base of the tag read no less than 0.5; d) a second maximum signal intensity of a specific base of the tag read or any feature with a Spearman correlation coefficient with the second maximum signal intensity of the specific base of the tag read no less than 0.

5.

7. The apparatus of claim 5 or 6, wherein, The apparatus further comprises a classifier determining unit, which comprises: a data acquiring subunit configured to acquire a training data set, the training data set being sequencing data of a plurality of samples with known assignment results of reads to samples of origins and true assignment results of the reads to the samples of origins, the training data set comprising a plurality of reads comprising at least a portion of a sequence to be sequenced of each sample and a specific tag sequence carried by the sample, the training data set comprising at least a specific number of reads; a representation subunit configured to divide the reads into two categories according to whether the assignment results and the true assignment results are consistent, and to represent each read in each category in an alternative form based on features of the read; a establishing subunit configured to count a number of reads in each category contained in each alternative form, to establish a correspondence between the alternative form and the category of the read; optionally, the features of the read comprise at least one of: a signal intensity of a portion of the insert read or any feature with a Spearman correlation coefficient with the signal intensity of the portion of the insert read no less than 0.5; a signal intensity of the tag read or any feature with a Spearman correlation coefficient with the signal intensity of the tag read no less than 0.5; a maximum signal intensity of a specific base of the tag read or any feature with a Spearman correlation coefficient with the maximum signal intensity of the specific base of the tag read no less than 0.5; a second maximum signal intensity of a specific base of the tag read or any feature with a Spearman correlation coefficient with the second maximum signal intensity of the specific base of the tag read no less than 0.

5. Optionally, the maximum signal intensity of a specific base of the tag readout or any feature with a Spearman correlation coefficient with the maximum signal intensity of the specific base of the tag readout not less than 0.5; and The second largest signal intensity of a specific base of the tag readout or any feature with a Spearman correlation coefficient with the second largest signal intensity of the specific base of the tag readout not less than 0.5, including: The maximum signal intensity of the first base with the first reciprocal Q value of the tag readout or any feature with a Spearman correlation coefficient with the maximum signal intensity of the first base with the first reciprocal Q value of the tag readout not less than 0.5; and The second largest signal intensity of the first base with the first reciprocal Q value of the tag readout or any feature with a Spearman correlation coefficient with the second largest signal intensity of the first base with the first reciprocal Q value of the tag readout not less than 0.5; Optionally, the maximum signal intensity of a specific base of the tag readout or any feature with a Spearman correlation coefficient with the maximum signal intensity of the specific base of the tag readout not less than 0.5; and The second largest signal intensity of a specific base of the tag readout or any feature with a Spearman correlation coefficient with the second largest signal intensity of the specific base of the tag readout not less than 0.5, including: The maximum signal intensity of the second base with the second reciprocal Q value of the tag readout or any feature with a Spearman correlation coefficient with the maximum signal intensity of the second base with the second reciprocal Q value of the tag readout not less than 0.5; and The second largest signal intensity of the second base with the second reciprocal Q value of the tag readout or any feature with a Spearman correlation coefficient with the second largest signal intensity of the second base with the second reciprocal Q value of the tag readout not less than 0.5; Optionally, the maximum signal intensity of a specific base of the tag readout or any feature with a Spearman correlation coefficient with the maximum signal intensity of the specific base of the tag readout not less than 0.5; and The second largest signal intensity of a specific base of the tag readout or any feature with a Spearman correlation coefficient with the second largest signal intensity of the specific base of the tag readout not less than 0.5, including: The maximum signal intensity of the third base with the third reciprocal Q value of the tag readout or any feature with a Spearman correlation coefficient with the maximum signal intensity of the third base with the third reciprocal Q value of the tag readout not less than 0.5; and The second largest signal intensity of the third base with the third reciprocal Q value of the tag readout or any feature with a Spearman correlation coefficient with the second largest signal intensity of the third base with the third reciprocal Q value of the tag readout not less than 0.

5.

8. The apparatus of claim 7, wherein, The representation sub-unit is used to: Normalize the feature value corresponding to the feature of the read, to obtain a normalized feature value; Quantize the normalized feature value into an alternative form. 9.A computer readable storage medium, for storing a program executed by a computer, and execution of the program comprises completing the multiple sample analysis method according to any one of the preceding 1-4. 10.A computing device, comprising: a memory, for storing a computer executable program; one or more processors, for executing the computer executable program and implementing the multiple sample analysis method according to any one of the preceding 1-4.

Citation Information

Patent Citations

  • Base recognition methods, systems and sequencing systems

    CN112823352B

  • Sequencing method

    CN113293205A

  • Method for constructing sequencing template based on image, and base recognition method and device

    WO2020037574A1

  • Method and device for cancer risk prediction

    CN116403644A

  • Sequencing data processing method and device, computing device and computer readable medium

    CN116721701A