Methods and Systems for Filtering Sequence Contamination Between Samples in High-Throughput Sequencing of Immune Library Repositories
By constructing a contamination risk prediction model and causal graph, and combining dynamic filtering and cross-batch background libraries, the problem of identifying and removing sequence contamination between samples in high-throughput sequencing of immune repertoires was solved, achieving high-precision contamination filtering and experimental optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHANGSHA WEISHI MEDICAL LAB CO LTD
- Filing Date
- 2026-01-28
- Publication Date
- 2026-04-21
AI Technical Summary
Existing high-throughput sequencing technologies for immune repertoires are ineffective at identifying and eliminating cross-contamination when dealing with sequence contamination between samples. In particular, during multi-well plate library construction, traditional methods cannot adapt to the differences in conditions between different experimental batches, resulting in a high false positive rate and a lack of ability to trace the source of contamination.
By combining a pollution risk prediction model with sample metadata, the pollution risk coefficient and noise baseline are dynamically calculated, a pollution causal graph is constructed, and precise filtering is performed through a three-stage filtering process and a multi-index fusion model. An updatable cross-batch background pollution database is also constructed to form an intelligent closed-loop system.
It significantly improves the accuracy and robustness of contaminated sequence filtering, can distinguish between real clones and contaminated sequences under different experimental conditions, systematically reduces contamination, and enhances data reliability and the ability to optimize experimental procedures.
Smart Images

Figure CN121570884B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical testing technology, and in particular to a method and system for filtering sequence contamination between samples in high-throughput sequencing of immune repertoires. Background Technology
[0002] High-throughput sequencing of immune repertoires aims to comprehensively assess the diversity of the immune system and identify and quantify proliferating clones of the immune response in vivo. This technology is extremely sensitive to sequence contamination because data analysis requires precision down to the clonal level; any contamination can directly affect the accuracy of clonal analysis. Especially when using multi-well plates for high-throughput automated library construction, effectively assessing and controlling sequence contamination between samples is a key challenge in ensuring data reliability.
[0003] Existing contamination treatment methods have significant limitations. Some technical solutions focus on isolating environmental contaminants through fully enclosed operating platforms, but this approach struggles to effectively address cross-contamination between samples. At the bioinformatics analysis level, existing contamination detection tools are often designed for whole-genome or whole-exome sequencing. When applied to sequencing data from immune repertoires with smaller target regions, errors can easily occur, potentially leading to increased false positive rates in panel sequencing that only covers a portion of genes. Furthermore, traditional methods rely on fixed frequency thresholds or static standards to filter contamination, failing to adapt to variations in sequencing depth, library concentration, and other conditions between different experimental batches. They lack the ability to trace contamination sources, and for systemic background contamination caused by reagents or the environment, analysis of a single batch is often insufficient to effectively identify and remove it.
[0004] This invention addresses the aforementioned problems by proposing a method and system for filtering sequence contamination between samples in high-throughput sequencing of immune repertoires. First, a contamination risk prediction model is used to dynamically calculate the inter-sample contamination risk coefficient and the overall batch noise baseline, combined with sample metadata, and a contamination causal graph is constructed to trace the contamination source. Based on this, a dynamically adjustable three-stage core filtering process is employed. Using reference sample calibration and dynamically setting the filtering threshold based on the risk coefficient, precise filtering of cross-contamination, abnormal clone propagation, and PCR or sequencing error products is achieved. After filtering, the system further uses a multi-index fusion model to perform sample-level contamination risk scoring and graded early warning, and utilizes an updatable cross-batch background contamination library for collaborative filtering of systemic background noise. Finally, the system feeds back newly identified contamination patterns to update the risk model and knowledge base, forming an intelligent closed-loop system with continuous learning and optimization capabilities. Summary of the Invention
[0005] To overcome the problems mentioned in the background art, the present invention proposes a method and system for filtering sequence contamination between samples in high-throughput sequencing of immune repertoires.
[0006] The technical solution of this invention is: a method for filtering sequence contamination between samples in high-throughput sequencing of immune repertoires, comprising:
[0007] S11: Preprocessing and contamination risk assessment: Receive sequencing data, sample metadata, and reference sample data; Based on the sample metadata, calculate the inter-sample contamination risk coefficient and the overall batch data noise baseline using a contamination risk prediction model, and construct a historical contamination causal graph.
[0008] S12: Dynamic core filtering, which uses reference sample data, contamination risk coefficient and data noise baseline to dynamically adjust the filtering threshold and standard, and performs three-stage sequence filtering on the sequencing data in sequence, including low-frequency filtering within lanes, sample ratio filtering between lanes and nucleotide sequence diversity verification.
[0009] S13: Pollution diagnosis and sample-level early warning. Based on the filtering results, calculate the multi-dimensional filtering indicators for each sample, and generate sample-level pollution risk scores and graded early warnings through the sample pollution risk model.
[0010] S14: Cross-batch collaborative filtering compares the filtered sequences with an updatable cross-project background contamination library and removes matching background contamination sequences;
[0011] S15: System feedback and update. The new pollution patterns and background clone sequences identified in this filtering are fed back and updated to the pollution risk prediction model, historical pollution causal diagram and cross-project background pollution database.
[0012] As a preferred option, the pretreatment and pollution risk assessment steps specifically include:
[0013] S21: Receive the input sequencing data, sample metadata, and sequencing data of a reference sample composed of known clones;
[0014] S22: The pollution risk prediction model trained based on historical data calculates the pollution risk coefficient between any two samples according to the input sample metadata, and calculates the data noise baseline representing the overall background noise level of the current batch of experiments.
[0015] S23: Based on historical batch data, using samples as nodes and the identified pollution directions as directed edges, construct a pollution causal graph that displays the pollution propagation path.
[0016] Preferably, in the pollution risk prediction model trained on historical data, when calculating the pollution risk coefficient between any two samples based on the input sample metadata, and when calculating the data noise baseline representing the overall background noise level of the current batch of experiments, the specific steps include:
[0017] S31: Feature normalization, which normalizes the numerical features in the input sample metadata, including sequencing depth, library concentration and PCR cycle number.
[0018] S32: Individual pollution risk score calculation. The normalized features and the spatial location information of the sample in the lane are input into the pollution risk prediction model. The pollution risk prediction model calculates the individual pollution risk score for each sample through pre-trained weights.
[0019] S33: Calculation of the contamination risk coefficient between samples. Based on the individual contamination risk scores of each pair of samples and their proximity relationship in the experimental operation, the contamination risk coefficient between any two samples is calculated.
[0020] S34: Overall data noise baseline calculation. Based on the individual contamination risk scores of all samples in the current batch, a comprehensive index representing the overall background noise level of this batch is calculated, namely the overall data noise baseline.
[0021] As a preferred method, when constructing a pollution causal graph showing the pollution propagation path based on historical batch data, using samples as nodes and identified pollution directions as directed edges, the specific steps include:
[0022] S41: Data input, receiving historical batch data, including clonal composition information for each historical sample, as well as potential contamination pairs and their directions identified from historical analysis;
[0023] S42: Graph structure definition, define a directed graph as a pollution causal graph;
[0024] S43: Contamination Relationship Identification: For any two historical samples, determine whether there is a potential contamination relationship between the two historical samples;
[0025] S44: Directed edge construction and weight assignment. For the identified contamination relationship, a directed edge is constructed and a weight is assigned to the constructed directed edge.
[0026] S45: Pollution source identification and analysis. Based on the constructed pollution causal graph, each vertex in the graph is analyzed, and its degree as a pollution source and victim is evaluated by comparing the total weight of its outgoing and incoming edges.
[0027] As a preferred option, low-frequency filtering within the swimming lane specifically includes:
[0028] S51: For all samples within the same sequencing lane, merge all their clone sequences and calculate the frequency of each clone in each sample;
[0029] S52: Analyze the data of the reference sample in the same lane and identify all known true clones; for each known clone, calculate the frequency ratio of its contamination to any other sample in the same lane, and dynamically determine the initial frequency ratio threshold based on the distribution of all calculated frequency ratio values for all known clones.
[0030] S53: Based on the calculated pollution risk coefficient, the initial threshold is adjusted to obtain the personalized final frequency ratio threshold used for this comparison;
[0031] S54: For any two samples within a lane and any clone they share, if their frequency ratio satisfies:
[0032] ;
[0033] Then, clone c in the low-frequency sample is determined to be caused by contamination from the high-frequency sample, and clone c is removed from the clone list of sample j. The frequency of clone c in sample i, Let c be the frequency of clone c in sample j.
[0034] As a preferred method, the inter-lane sample ratio filtering is specifically as follows:
[0035] S61: For each clone sequence that still exists after filtering within the lane, calculate the proportion of its occurrence in each sequencing lane.
[0036] S62: Based on the calculated data noise baseline representing the overall background noise level of the current batch, the exceedance multiple threshold for judging the proportion of abnormality is dynamically determined through a preset mapping function.
[0037] S63: Contamination clone determination. For a specific clone and lane, determine whether it is an abnormally spread contaminant in that lane by comparing the relationship between its appearance ratio in the sample in that lane and the average appearance ratio in the sample in other lanes.
[0038] S64: For a clone that is determined to be contaminated in the lane, the clone is sorted from low to high frequency among all samples in which the clone appears. In this order, the clone sequence is removed from the samples with the lowest frequency. After each removal, the sample occurrence ratio of the clone in the lane is recalculated until the sample occurrence ratio no longer meets the contamination determination criteria.
[0039] As a preferred option, pollution diagnosis and sample-level early warning specifically include:
[0040] S71: For each sample, based on the filtering results, calculate the sequence loss rate, clone overflow, and filtered clone characteristic indicators in the three stages of low-frequency filtering within lanes, sample ratio filtering between lanes, and nucleotide sequence diversity verification.
[0041] S72: Input the calculated multi-dimensional indicators into the sample pollution risk model, and calculate the comprehensive pollution risk score for each sample through a preset weighted fusion algorithm;
[0042] S73: Based on the calculated comprehensive pollution risk score of the samples, the samples are divided into different pollution risk levels, and corresponding early warning information is generated.
[0043] As a preferred embodiment, cross-batch collaborative filtering specifically includes:
[0044] S81: Based on historical batch data, construct and continuously update a cross-project background contamination library. The background contamination library records low-frequency cloning sequences and their meta-information that repeatedly appear in negative controls and irrelevant samples.
[0045] S82: Compare all the clone sequences retained after filtering in the current batch with the clone sequences recorded in the background contamination library;
[0046] S83: If a clone sequence in the current batch matches a sequence in the background contamination library, the clone is determined to be a systematic background contamination and is removed from the data in the current batch.
[0047] As a preferred option, system feedback and updates specifically include:
[0048] S91: Add the newly obtained sample-to-pollution relationship data and its corresponding sample metadata to the historical training dataset, and use the updated dataset to periodically retrain the pollution risk prediction model.
[0049] S92: The identified contamination relationships between samples are added as new directed edges and their weights to the historical contamination causal graph;
[0050] S93: After evaluating residual clones in samples marked as high risk, as well as clone sequences that are highly similar to but not completely matched with the current background contamination library, according to preset rules, the clones that meet the conditions will be added as new background contamination sequences to the cross-project background contamination library.
[0051] An inter-sample sequence contamination filtering system for high-throughput sequencing of immune repertoires includes:
[0052] The data receiving and preprocessing module is used to receive sequencing data, sample metadata, and reference sample data.
[0053] The pollution risk assessment module is used to calculate the pollution risk coefficient between samples and the overall data noise baseline of the batch based on sample metadata through a pollution risk prediction model, and to construct and update historical pollution causal graphs.
[0054] The dynamic core filtering module is used to dynamically adjust the filtering threshold and standard using reference sample data, contamination risk coefficient and data noise baseline, and sequentially perform low-frequency filtering within lanes, sample ratio filtering between lanes and nucleotide sequence diversity verification on the sequencing data.
[0055] The pollution diagnosis and early warning module is used to calculate multi-dimensional filtering indicators for each sample based on the filtering results, and generate sample-level pollution risk scores and graded early warnings through the sample pollution risk model.
[0056] The cross-batch collaborative filtering module is used to compare the filtered sequences with the background contamination library and remove matching background contamination sequences.
[0057] The feedback and update module is used to feed back and update the identified new pollution patterns and background clone sequences to the pollution risk prediction model, historical pollution causal diagram and background pollution database.
[0058] The beneficial effects of this invention are:
[0059] 1. Compared to existing technologies that typically use fixed frequency ratio thresholds or static standards to filter cross-contamination between samples, this method cannot adapt to changes in different experimental batches and sample characteristics, and may lead to over-filtering of real signals or residual contamination. This invention adopts a strategy of dynamically adjusting the filtering threshold based on a contamination risk prediction model. It calculates the contamination risk coefficient and data noise baseline through sample metadata, and realizes personalized threshold setting. This dynamic adaptive mechanism significantly improves the filtering accuracy and can intelligently distinguish between real clones and contaminated sequences under different experimental conditions, effectively balancing sensitivity and specificity.
[0060] 2. Compared to existing technologies that only focus on sequence filtering and lack the ability to track and visualize pollution sources, resulting in a lack of data support for optimizing experimental procedures, this invention introduces a pollution causal graph construction technique. This technique constructs historical pollution relationships into a directed graph structure and locates the main pollution sources and affected nodes by calculating the net pollution contribution value of the samples. This design not only visualizes the pollution propagation path but also provides clear guidance for systematically reducing pollution at the experimental operation level, achieving a leap from simple filtering to in-depth diagnosis.
[0061] 3. Compared to existing technologies that mainly focus on filtering contamination within a single batch and lack effective means for systematic background contamination across batches and projects, making it difficult to completely remove some stubborn contaminants, this invention constructs an updatable cross-project background contamination database. By integrating historical batch data, it identifies low-frequency clones that repeatedly appear in negative controls and irrelevant samples, and uses this as a basis for collaborative filtering. This method breaks through the limitations of single-batch analysis, forming a continuously accumulating contamination fingerprint database, and has a strong ability to identify and remove systematic contamination from reagents, the environment, etc.
[0062] 4. Compared to existing technologies where each filtering step is relatively independent and lacks the ability to learn and evolve from new data, resulting in stagnant system performance, this invention designs a complete closed-loop feedback and update mechanism. It uses the new pollution relationships and sample metadata confirmed by each analysis to periodically retrain the pollution risk prediction model and update the pollution causal graph and background pollution database. This enables the system to have the ability to learn and continuously optimize itself, and the filtering performance continues to improve over time, forming an intelligent ecosystem that becomes more accurate with use.
[0063] 5. Compared with existing technologies that rely on a single principle or a single perspective for filtration, making it difficult to comprehensively cover multiple types of contamination and creating detection blind spots, this invention designs a three-stage progressive filtration process, which sequentially performs low-frequency filtration within lanes, sample ratio filtration between lanes, and nucleotide sequence diversity verification. It combines multiple mechanisms such as in-laboratory correction, statistical anomaly detection, and biochemical principle verification. This multi-level collaborative verification strategy achieves comprehensive coverage of contamination from different sources, significantly improving the robustness and reliability of the method. Attached Figure Description
[0064] Figure 1 The diagram shows a flowchart of the method for filtering sequence contamination between samples in high-throughput sequencing of immune repertoires according to the present invention.
[0065] Figure 2 The diagram shown is a schematic representation of the structure of the high-throughput sequencing sample sequence contamination filtering system of the immune repertoire of the present invention. Detailed Implementation
[0066] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0067] Please see Figure 1 This invention provides an embodiment: a method for filtering sequence contamination between samples in high-throughput sequencing of immune repertoires, comprising:
[0068] Step 1: Pretreatment and Pollution Risk Assessment
[0069] Receive sequencing data, sample metadata, and reference sample data; based on the sample metadata, calculate the inter-sample contamination risk coefficient and the overall batch data noise baseline using a contamination risk prediction model, and construct a historical contamination causality graph, specifically including:
[0070] It receives input sequencing data, sample metadata, and sequencing data of a reference sample composed of known clones. The sample metadata includes the sequencing depth, library concentration, PCR cycle number, and spatial position information of the sample in the sequencing lane.
[0071] The pollution risk prediction model trained based on historical data calculates the pollution risk coefficient between any two samples based on the input sample metadata, and calculates the data noise baseline representing the overall background noise level of the current batch of experiments.
[0072] Based on historical batch data, a pollution causal graph is constructed, using samples as nodes and the identified pollution directions as directed edges, to display the pollution propagation path.
[0073] In this embodiment, when calculating the pollution risk prediction model trained based on historical data, and calculating the pollution risk coefficient between any two samples according to the input sample metadata, and when calculating the data noise baseline representing the overall background noise level of the current batch of experiments, the specific steps include:
[0074] Feature normalization normalizes the numerical features in the input sample metadata, including sequencing depth, library concentration, and PCR cycle number, so that they fall into a uniform numerical range.
[0075] The individual pollution risk score is calculated by inputting normalized features and the sample's spatial location information in the swimlane into the pollution risk prediction model. The pollution risk prediction model calculates an individual pollution risk score for each sample using pre-trained weights. The calculation formula is as follows: ;in, The pollution risk score representing sample i. This represents the normalized sequencing depth of sample i. The normalized library concentration of sample i, The normalized PCR cycle number characteristic value of sample i. The spatial risk factor is determined based on the position of sample i in the swimlane. , , and The weight coefficients of each feature are determined by the pollution risk model through learning from historical pollution data;
[0076] The inter-sample contamination risk coefficient is calculated based on the individual contamination risk scores of each pair of samples and their proximity in the experimental procedure. The formula is as follows: ;in, The risk coefficient representing the contamination of sample i into sample j. For sample i, the individual pollution score is... For sample j, the individual pollution score, The distance between sample i and sample j in the experimental procedure can be quantified as their physical spatial distance on the PCR plate, whether they belong to the same mixing chamber, and whether they are adjacent sample loading orders. It is a decreasing function of the operation distance. The closer the operation distance between two samples, the larger the function value, and vice versa.
[0077] The overall data noise baseline is calculated based on the individual contamination risk scores of all samples in the current batch. It is a comprehensive index characterizing the overall background noise level of this batch, i.e., the overall data noise baseline. The calculation formula is as follows: ;in, This represents the total number of samples in the current batch. This is the set of individual pollution risk scores for all samples. This indicates the calculation of the p-th percentile of the set, where p is a parameter determined by optimization based on historical data.
[0078] In this embodiment, when constructing a pollution causal graph that displays the pollution propagation path based on historical batch data, using samples as nodes and identified pollution directions as directed edges, the specific steps include:
[0079] Data input receives historical batch data, including clonal composition information for each historical sample, as well as potential contamination pairs and their directions identified from historical analysis.
[0080] Graph structure definition: Define a directed graph. As a pollution causal graph, where: vertex set Each vertex in Represents an independent sample; directed edge set Each directed edge in This represents the contamination propagation relationship from sample i to sample j;
[0081] Contamination relationship identification: For any two historical samples, determine whether there is a potential contamination relationship between the two historical samples;
[0082] Directed edge construction and weight assignment: For each identified contamination relationship, a directed edge is constructed, and a weight is assigned to the constructed directed edge. The weight is calculated based on the frequency, quantity, and confidence level of the contamination clones. The calculation formula is as follows: ;in, The assigned weights are used to quantify the contamination intensity from sample i to sample j. Let J be the set of all clones from sample i that were determined to be contaminated to sample j. To clone the frequency of c in the source sample i, The frequency of clone c in source sample j;
[0083] Pollution source identification and analysis, based on a constructed pollution causal graph, analyzes each vertex in the graph, and assesses its degree as a pollution source and victim by comparing the total weight of its outgoing and incoming edges. The calculation formula is: Net Pollution Contribution. ;in, The net pollution contribution value for sample i. This represents the weight of the directed contamination edge from sample i to sample k. This represents the total contamination output of sample i for all other samples. The weights of the directed contamination edges from sample m to sample i. This represents the total contamination input that sample i receives from all other samples.
[0084] In this embodiment, when identifying contamination relationships, the following rules are used to determine whether there is a potential contamination relationship between two historical samples:
[0085] For any two historical samples i and j, if the following condition is met, then it is determined that there exists a potential contamination relationship from sample i to sample j:
[0086] Condition: The frequency of clone c in sample i Significantly higher than the frequency in sample j And the frequency is higher than Exceeding a threshold set based on historical data or experimental conditions Furthermore, clone c was identified as a true clone rather than a contaminant in sample i.
[0087] In this embodiment, if the net pollution contribution of a sample is a significant positive value, the sample is determined to be the main pollution source node; if it is a significant negative value, it is determined to be the main pollution victim node.
[0088] Specifically, pollution causal diagrams are used to trace the source of pollution and visualize the transmission path of pollution between samples. The analysis results are used to guide the optimization of experimental procedures.
[0089] In this embodiment, the present invention first uses a contamination risk prediction model, combined with multidimensional information such as sequencing depth, library concentration, PCR cycle number, and spatial location of the samples, to calculate the contamination risk coefficient between samples and the batch noise baseline, and constructs a contamination causal graph to trace the source of contamination. Based on this, a dynamically adjustable three-stage core filtering process is adopted: the first stage uses reference samples for calibration and combines the risk coefficient to dynamically set the clone frequency ratio threshold within the lane, accurately removing cross-contamination; the second stage dynamically adjusts the judgment criteria for the proportion of samples appearing between lanes based on the batch noise baseline, identifying and eliminating abnormally propagated clones; the third stage removes possible PCR or sequencing error products by verifying the nucleotide coding diversity of the amino acid sequence. After filtering, the system further uses a multi-index fusion model to score the contamination risk of each sample and provide graded early warnings, and uses an updatable cross-batch background contamination library for final collaborative filtering to remove systemic background noise from reagents, the environment, etc. Finally, the system feeds back the newly identified contamination patterns to update the risk model, causal graph, and background library, forming a self-optimizing closed loop. This approach systematically addresses the challenges of traditional methods that rely on fixed thresholds, lack adaptive capabilities, and fail to trace contamination sources. Through multi-dimensional verification, dynamic threshold adjustment, machine learning prediction, and closed-loop learning mechanisms, it significantly improves the accuracy and reliability of immune repertoire data in applications such as tumor immunology and vaccine evaluation, and can also guide the optimization of experimental procedures.
[0090] Step 2: Dynamic Core Filtering
[0091] Using reference sample data, contamination risk coefficient, and data noise baseline, the filtering threshold and standard are dynamically adjusted to perform three-stage sequence filtering on the sequencing data, including low-frequency filtering within lanes, sample ratio filtering between lanes, and nucleotide sequence diversity verification.
[0092] In this embodiment, the low-frequency filtering within the swimming lane specifically involves:
[0093] For all samples within the same sequencing lane, merge all their clone sequences and calculate the frequency of each clone in each sample. Specifically, the frequency of clone c in sample i is calculated. The calculation formula is as follows: ;in, To count the original sequence of clone c in sample i, The total sequence count for all clones in sample i;
[0094] Analyze the data from the reference samples in the same lane to identify all known true clones; for each known clone, calculate the frequency ratio of its contamination to any other sample in the same lane. Based on the distribution of all frequency ratio values calculated from all known clone pairs, an initial frequency ratio threshold is dynamically determined, where the frequency ratio... The calculation formula is: ;in, For a single pollution frequency ratio observation, Given the frequency of known clones in the reference sample, Let be its frequency in sample s; where, when dynamically determining the initial frequency-to-ratio threshold, specifically, a specified high quantile of the distribution is calculated. ,Right now: ;in, The initial frequency ratio threshold is calculated based on the reference sample. This indicates the calculation of the p-quantile, where p is a parameter preset based on experience. For a single known clone in the reference sample, For any sample in the same lane other than the reference sample itself;
[0095] Based on the calculated pollution risk coefficient, the initial threshold is adjusted to obtain the personalized final frequency ratio threshold used for this comparison. The calculation expression is as follows: ;in, To personalize the final frequency ratio threshold, Let be the pollution risk coefficient between sample i and sample j. It is a monotonically decreasing adjustment function;
[0096] For any two samples within a lane and any clone they share, if their frequency ratio satisfies: Then, clone c in the low-frequency sample is determined to be caused by contamination from the high-frequency sample, and clone c is removed from the clone list of sample j. The frequency of clone c in sample i, Let c be the frequency of clone c in sample j.
[0097] In this embodiment, the inter-lane sample ratio filtering is specifically as follows:
[0098] For each clone sequence that survives the in-lane filtering, the proportion of its occurrence in each sequencing lane is calculated using the following formula: ;in, For each sequencing lane, clone sequence c is used. The proportion of samples appearing in the sample, For in the swimming lane The number of samples that detected clone c was... For this lane The total number of samples included;
[0099] Based on the calculated data noise baseline representing the overall background noise level of the current batch, a threshold for exceeding the standard multiple used to determine the proportion of abnormalities is dynamically determined through a preset mapping function. The calculation expression of the mapping function is as follows: ;in, The threshold for exceeding the limit is dynamically adjusted. The preset threshold benchmark value, This is the preset strictness adjustment coefficient. This represents the baseline value for data noise.
[0100] Contamination clone determination: For a specific clone and lane, determine whether it is an abnormally spread contaminant in that lane by comparing the proportion of its appearance in the sample in that lane with the average proportion of its appearance in the sample in other lanes to see if it exceeds the exceedance multiple threshold.
[0101] For clones identified as contaminating the lane, the clone is sorted from lowest to highest frequency among all samples in which it appears. Then, the clone sequence is removed from the samples with the lowest frequency in that order. After each removal, the proportion of the clone in the lane is recalculated until the proportion no longer meets the contamination criteria.
[0102] In this embodiment, the specific steps for determining contaminated clones are as follows:
[0103] For a clone c to be evaluated, and its specific lane... Calculate the average proportion of samples where this clone appears in all other lanes using the following formula:
[0104] ;
[0105] in, The average proportion of clone c appearing in samples across all other lanes. This is the set that includes all other lanes except the one currently being analyzed. Indicates the number of other lanes. The proportion of clone c appearing in lane l;
[0106] The clone is in the swimming lane If the proportion of samples appearing in a lane satisfies the following formula, then the clone is determined to be in the lane. The contamination is spreading abnormally.
[0107] ;
[0108] in, For cloning c in the swim lane The proportion of samples appearing in the sample, The threshold for exceeding the limit is dynamically adjusted. The average proportion of clone c appearing in all other lanes.
[0109] In this embodiment, the nucleotide sequence diversity verification specifically involves:
[0110] The remaining cloned nucleotide sequences after filtering are translated into their corresponding amino acid sequences; for the same amino acid sequence, it is checked whether it is encoded by multiple different nucleotide sequences; if an amino acid sequence is encoded by only one nucleotide sequence, the nucleotide sequence is determined to be a possible amplification or sequencing error product and is removed.
[0111] In this embodiment, a three-stage progressive, threshold-adaptive sequence filtering strategy is employed in the dynamic core filtering stage to accurately distinguish between real clones and contamination signals. First, in low-frequency filtering within lanes, an initial frequency ratio threshold is determined using the actual contamination data distribution of a reference sample. This threshold is then individually tightened using the inter-sample contamination risk coefficient calculated in step one, forming a dynamic threshold. This threshold judges the frequency ratio of shared clones in any two samples within the same lane, efficiently removing low-frequency "shadow" clones caused by cross-contamination from adjacent samples. Second, in inter-lane sample proportion filtering, an abnormal proportion threshold is dynamically set based on the data noise baseline (DNB), which represents the overall batch noise level. This identifies and removes clones with abnormally high proportions in specific lanes. These clones typically indicate uneven contamination across lanes that occurred during PCR amplification or pooling. Finally, nucleotide sequence diversity verification is performed. After translating residual sequences into amino acids, amino acid sequences encoded by a single nucleotide sequence are removed. These sequences are highly likely to be erroneous products generated during PCR or sequencing, rather than genuine biological clones. This technical solution upgrades fixed empirical thresholds to intelligent thresholds that are dynamically adjusted according to experimental quality and sample characteristics through a multi-level filtering logic of "internal reference calibration - personalized risk adjustment - cross-lane mode verification - biochemical principle validation". It achieves comprehensive coverage from experimental contamination to technical errors, thereby significantly improving the accuracy and robustness of contamination sequence removal without excessive loss of real low-frequency clones, laying a reliable data foundation for subsequent high-sensitivity immune repertoire analysis.
[0112] Step 3: Pollution Diagnosis and Sample-Level Early Warning
[0113] Based on the filtering results, multi-dimensional filtering indicators are calculated for each sample. A sample-level pollution risk score and graded early warning are generated using a sample pollution risk model, specifically including:
[0114] For each sample, based on the filtering results, the sequence loss rate, clone overflow, and filtered clone characteristic indicators are calculated in three stages: low-frequency filtering within lanes, sample ratio filtering between lanes, and nucleotide sequence diversity verification.
[0115] The calculated multi-dimensional indicators are input into the sample pollution risk model. A pre-defined weighted fusion algorithm is used to calculate the comprehensive pollution risk score for each sample. The calculation formula is as follows: ;in, The comprehensive pollution risk score for sample i. This indicates that the indicator has been normalized. This represents the total loss rate of sample i. The clone overflow index for sample i. Let i be the filtered clone feature index. , and These are the preset weighting coefficients for the corresponding indicators;
[0116] Based on the calculated comprehensive pollution risk score of the samples, the samples are divided into different pollution risk levels and corresponding early warning information is generated. Specifically, the comprehensive pollution risk score of the samples is compared with a preset threshold range to classify the risk level.
[0117] In this embodiment, the calculation of the sequence loss rate index, clone overflow index, and filtered clone feature index is as follows:
[0118] Calculate the sequence loss rate for sample i, which includes the total loss rate and the stage-specific loss rate:
[0119] Total loss rate: ;
[0120] in, Let i be the total loss rate of sample i. Let be the total number of sequences filtered by sample i in the three stages of core filtering. Let be the original total number of sequences for sample i;
[0121] The clone spillover index for sample i is calculated to quantify the degree of high-frequency clone contamination in sample i compared to other samples in the lane. The calculation formula is as follows:
[0122] ;
[0123] in, The clone overflow index for sample i. Let i be the set of the top K most frequent clones in sample i. This is an indicator function; if a high-frequency clone c in sample i is determined to have contaminated sample j in the same lane, its value is 1; otherwise, it is 0. The frequency of clone c in sample i, This is the frequency reference value used for normalization;
[0124] The filtered clone feature index of sample i is calculated to reflect the overall distribution characteristics of the filtered clones. The calculation formula is as follows:
[0125] ;
[0126] in, For the filtered clone feature index of i, Let M be the set of the original frequencies of all M clones filtered out from sample i in sample i. This indicates the calculation of the geometric mean of the set.
[0127] In this embodiment, the present invention constructs a multi-dimensional and quantitative sample-level quality control and intelligent decision-making mechanism in the pollution diagnosis and sample-level early warning stages. This technical solution is no longer limited to sequence filtering itself, but rather integrates detailed information from the core filtering process to calculate three key diagnostic indicators for each sample: the total sequence loss rate directly reflects the proportion of filtered sequences in the sample; the clone overflow indicator specifically quantifies the "source attribute" intensity of high-frequency clone contamination in the sample compared to other samples in the lane; and the filtered clone characteristic indicator uses the geometric mean to characterize the overall frequency distribution of removed clones. After normalization, these indicators are used to calculate a comprehensive pollution risk score through a preset weighted fusion model, and based on this score, samples are automatically classified into "low risk," "medium risk / early warning," and "high risk / recommendation for removal" levels. This solution represents a leap from "sequence filtering" to "sample diagnosis," with the following benefits: it makes the contamination information that was originally hidden in the filtering process explicit and quantifiable, providing researchers with a clear sample quality assessment report; its tiered early warning mechanism can automatically identify severely contaminated samples that may need to be reviewed or removed, avoiding misjudgments caused by a single indicator (such as total loss rate), significantly improving the precision and intelligence of data quality control, and providing direct data support for tracing the root cause of contamination and optimizing experimental procedures.
[0128] Step 4: Cross-batch collaborative filtering
[0129] The filtered sequences are compared with an updatable cross-project background contamination library, and matching background contamination sequences are removed, specifically including:
[0130] Based on historical batch data, a cross-project background contamination library is constructed and continuously updated. The background contamination library records low-frequency cloning sequences and their metadata that repeatedly appear in negative controls and irrelevant samples.
[0131] Compare all the clone sequences retained after filtering in the current batch with the clone sequences recorded in the background contamination library;
[0132] If a clone sequence in the current batch matches a sequence in the background contamination library, the clone is determined to be a systematic background contamination and is removed from the data in the current batch.
[0133] In this embodiment, the construction and maintenance of the background contamination database specifically includes:
[0134] a. Summarize all historical batches of residual clone sequence data after core filtering, and identify candidate clones that meet the background contamination characteristics. The identification criteria are: the clone appears in at least M different batches of negative control and irrelevant samples, and its highest frequency in any single sample does not exceed the preset extremely low frequency threshold.
[0135] b. For each identified candidate background clone c, calculate its cross-batch integration frequency index, calculated using the following formula:
[0136] ;
[0137] in, To integrate frequency indicators, This is the set of all batches in which this clone has appeared. Let b be the set of all samples in batch b. The frequency of clone c in sample s;
[0138] c. Integrate frequency indicators Greater than the preset threshold And the number of batches is greater than or equal to the preset threshold. The clones are incorporated into a cross-project background contamination database; this database is automatically updated periodically to include new batches of data.
[0139] In this embodiment, the threshold condition for identifying candidate background clones is implemented through the following formulaic rules:
[0140] A clone c is identified as a candidate background clone if and only if both of the following conditions are met:
[0141] Condition 1: ;
[0142] Condition 2: ;
[0143] in, M represents the total number of batches in which clone c appears across different historical batches, where M is the preset minimum number of batches. This represents the highest frequency of this clone across all historical samples. The preset extremely low frequency threshold ranges from 0.0001% to 0.001%.
[0144] Specifically, background sequence alignment employs a sequence exact matching algorithm, and background contamination removal satisfies the following conditions:
[0145] If the similarity between the current batch of cloned sequences and the sequences in the background contamination library reaches a preset similarity threshold... (For exact matches, If the percentage of matches is 100%, it is considered a match; for a matching clone, regardless of its frequency in the current batch, it is removed from the sample.
[0146] In this embodiment, the present invention introduces a background contamination knowledge base system with self-learning and evolution capabilities during the cross-batch collaborative filtering stage to systematically remove persistent background noise that is difficult to identify in a single analysis. The core of this technical solution lies in constructing and dynamically updating a "cross-project background contamination database," which intelligently filters out clonal sequences that repeatedly appear in multiple different batches of negative controls or irrelevant samples but have extremely low frequencies in any single sample by integrating all historical batch data. These sequences are typically background noise from reagent components, environmental microorganisms, or laboratory-specific sources. During construction, the system employs a dual-threshold rule (the number of batches appearing M and the highest frequency threshold f_bg) to ensure the typical background characteristics of the sequences entering the database, and quantifies their contamination intensity by integrating frequency indicators. When processing new data, the residual sequences filtered through the aforementioned steps are precisely compared with this background database. Any matching sequence, regardless of its frequency in the current batch, is determined to be systematic background contamination and ultimately removed. The beneficial effects of this approach lie in its ability to overcome the limitations of single-batch analysis. By accumulating and sharing contamination "fingerprint" information across batches and projects, it achieves efficient identification and removal of hidden and persistent background contamination. This not only significantly improves data purity, particularly facilitating the detection of extremely low-frequency real biological signals (such as trace residual lesions and rare antigen-specific clones), but also establishes a continuously optimized quality control system through a shareable and updatable knowledge base, significantly enhancing the consistency and comparability of long-term laboratory data.
[0147] Step 5: System Feedback and Updates
[0148] The new pollution patterns and background clone sequences identified in this filtering process will be fed back and updated to the pollution risk prediction model, historical pollution causal diagram, and cross-project background pollution database, specifically including:
[0149] The newly obtained sample-to-pollution relationship data and their corresponding sample metadata are added to the historical training dataset, and the pollution risk prediction model is periodically retrained using the updated dataset.
[0150] The identified contamination relationships between samples are used as new directed edges and their weights, and then added to the historical contamination causal graph.
[0151] Residual clones within samples marked as high risk, as well as clone sequences that are highly similar to but not completely matched with the current background contamination library, are evaluated according to preset rules. Clones that meet the criteria are then added as new background contamination sequences to the cross-project background contamination library.
[0152] In this embodiment, when adding the newly obtained sample-to-pollution relationship data and its corresponding sample metadata to the historical training dataset, and periodically retraining the pollution risk prediction model using the updated dataset, the specific steps include:
[0153] Collect the pollutant relationship pairs between samples that were identified and confirmed in this analysis and that have a high confidence level according to statistical analysis, and form new pollutant relationship records. Each record includes: source sample and target sample identifiers, sample element information, and the calculated pollutant risk coefficient.
[0154] Add the newly added pollution relationship records and their corresponding sample metadata to the historical training dataset;
[0155] When the amount of newly added data reaches a preset threshold or at fixed time intervals, the model retraining process is triggered. The pollution risk prediction model is retrained using the updated historical training dataset to update the feature weight coefficients within the model.
[0156] In this embodiment, when the identified contamination relationships between samples are added to the historical contamination causal graph as new directed edges and their weights, the specific steps include:
[0157] Extract the contamination relationships between samples identified in this analysis. For each contamination relationship that meets the confidence level, locate the corresponding source sample vertex and target sample vertex in the vertex set of the historical contamination causal graph.
[0158] If a directed edge already exists between the source sample vertex and the target sample vertex, the weight of the edge is updated according to the pollution intensity obtained in this analysis and a preset strategy; if it does not exist, a new directed edge is created in the historical pollution causal graph and its weight is initialized according to the results of this analysis.
[0159] After the historical pollution causal graph is updated, the net pollution contribution of all relevant samples is automatically recalculated, and new major pollution source nodes and pollution victim nodes are marked.
[0160] In this embodiment, after evaluating residual clones in samples marked as high-risk, and clone sequences that are highly similar to but not completely matched with the current background contamination library according to preset rules, the process of adding eligible clones as new background contamination sequences to the cross-project background contamination library specifically includes:
[0161] Two types of clone sequences were collected as candidate background sequences in this analysis: the first type is residual clones with a frequency below a preset threshold in samples marked as high risk; the second type is clones whose similarity to sequences in the existing background contamination library exceeds the threshold but does not reach a complete match during the comparison.
[0162] The candidate background sequences were evaluated, and sequences that appeared in at least N independent negative controls or low-risk samples from this batch were selected.
[0163] The candidate background sequences for the evaluation criteria, along with the batch and frequency information of the samples in which they appear, are added as new entries to the cross-project background contamination database, and their entry time and source batch are recorded.
[0164] In this embodiment, the output of a single data analysis is used as a feedback signal to systematically update and enhance the system's three core knowledge bases: First, the high-confidence pollution relationships and their experimental metadata are added to the training set, and the pollution risk prediction model is periodically retrained, making its assessment of pollution risk between samples increasingly accurate with data accumulation; second, newly discovered pollution paths are used as directed edges to update the historical pollution causal graph in real time, continuously enriching the topology of the pollution propagation network, making pollution source tracing more comprehensive and visualization clearer; finally, potential background sequences that meet the characteristics in the current data (such as low-frequency residual clones in high-risk samples, sequences highly similar to the existing background database) are intelligently screened, and after evaluation, the cross-project background pollution database is expanded, making its coverage of "pollution fingerprints" increasingly complete. The beneficial effect of this scheme is that it transforms a one-time analysis process into an intelligent system with "memory" and "evolution" capabilities. Through iterative learning, the system's filtering accuracy, pollution source tracing ability, and background noise recognition ability can all continuously improve over time. This not only ensures the high quality and consistency of long-term data analysis results, but also allows data insights to guide the optimization of experimental processes (such as identifying easily contaminated links), truly realizing an intelligent closed loop from "data quality control" to "process improvement".
[0165] like Figure 2 As shown, this embodiment also provides a sequence contamination filtering system between samples in high-throughput sequencing of immune repertoires, including:
[0166] The data receiving and preprocessing module is used to receive sequencing data, sample metadata, and reference sample data.
[0167] The pollution risk assessment module is used to calculate the pollution risk coefficient between samples and the overall data noise baseline of the batch based on sample metadata through a pollution risk prediction model, and to construct and update historical pollution causal graphs.
[0168] The dynamic core filtering module is used to dynamically adjust the filtering threshold and standard using reference sample data, contamination risk coefficient and data noise baseline, and sequentially perform low-frequency filtering within lanes, sample ratio filtering between lanes and nucleotide sequence diversity verification on the sequencing data.
[0169] The pollution diagnosis and early warning module is used to calculate multi-dimensional filtering indicators for each sample based on the filtering results, and generate sample-level pollution risk scores and graded early warnings through the sample pollution risk model.
[0170] The cross-batch collaborative filtering module is used to compare the filtered sequences with the background contamination library and remove matching background contamination sequences.
[0171] The feedback and update module is used to feed back and update the identified new pollution patterns and background clone sequences to the pollution risk prediction model, historical pollution causal diagram and background pollution database.
[0172] Example 1: Application in the monitoring of minimal residual disease (MRD) in tumors
[0173] This embodiment simulates a minimal residual disease (MRD) monitoring scenario in leukemia patients after treatment. This scenario demands extremely high data purity because it requires identifying very low-frequency true tumor clone signals from a large amount of background signal; any residual contamination could lead to misdiagnosis of the condition.
[0174] First, the system receives sequencing data and metadata, including patient bone marrow samples, negative control samples, and known tumor clone sequences (as reference samples). The contamination risk prediction model focuses on calculating individual contamination risk scores for samples with extremely high sequencing depth and a high number of PCR cycles (due to the capture of low-frequency clones). Because the samples are located close together on the PCR plate, the model calculates a high inter-sample contamination risk coefficient between these clinical samples and the negative controls. Simultaneously, the baseline data noise calculated based on all samples in the current batch is at a high level, reflecting the inherent high background noise of MRD testing. Historical contamination causality plots show that samples with high tumor clone abundance have been repeatedly identified as contamination sources in similar previous testing batches.
[0175] Low-frequency filtering within lanes: The system calibrates thresholds using reference samples (known tumor clones). For example, in 95% of contamination events, the frequency ratio of clones in the reference sample to the contaminated sample exceeded 1000 times. Based on this, and combined with the high contamination risk coefficient calculated in step one, the system sets a very strict, personalized frequency ratio threshold for each pair of adjacent samples. Any low-frequency clonal sequences below this threshold, suspected of originating from a high-abundance sample and migrating to a low-abundance or negative sample, are precisely removed.
[0176] Inter-lane sample proportion filtering: The system detected an abnormally high proportion of a clone in one sequencing lane (e.g., 80%), while the average proportion in other lanes was only 5%. Due to the high baseline noise in this batch of data, the system dynamically set a low fold-over threshold, and the abnormal distribution pattern of this clone triggered the filtering condition. The system then removed the clone sequentially from the lowest frequency samples in that lane until its distribution returned to normal. This eliminated cross-lane contamination that might have been caused by improper pooling operations.
[0177] Nucleotide sequence diversity verification: Finally, the system translated the remaining clones into amino acid sequences and found that a few amino acid sequences were encoded by unique nucleotide sequences. These sequences were determined to be more likely to originate from PCR or sequencing errors rather than true biological clones and were therefore discarded.
[0178] After filtering, the system calculated the indicators for each sample. One negative control sample showed a high total sequence loss rate and filtered clone characteristic indicators (low original frequency of filtered clones), indicating contaminant intrusion. The system marked it as "high risk" and recommended that the laboratory review the experimental procedure for this control sample. The primary patient sample, on the other hand, showed low clone spillover indicators, indicating that it was not the primary source of contamination. After comprehensive scoring, it was rated as "low risk," and the data was considered reliable.
[0179] The filtered data was compared with a cross-project background contamination library. The system identified some sequences that perfectly matched background clone sequences recorded in the library that are commonly found in laboratory reagents or the environment. Although these sequences were not frequent in this sample, they were removed as known systematic noise, further ensuring data purity and clearing the way for the detection of true extremely low frequency MRD signals.
[0180] The contamination relationships identified in this analysis (such as contamination pathways from a patient sample to a negative control) and their metadata were added to the training dataset for subsequent periodic retraining of the contamination risk prediction model. New contamination pathways were added as directed edges to the historical contamination causal graph, enriching the knowledge of laboratory contamination transmission patterns. New low-frequency residual clones discovered in the negative control, after being assessed as meeting background contamination characteristics, were added to the cross-project background contamination database, enabling the system to identify such contamination earlier in the future.
[0181] Example 2: Application in Vaccine Immunological Response Assessment
[0182] This embodiment simulates the application of evaluating the dynamic changes of a volunteer's immune repertoire in a clinical trial of vaccination. This scenario focuses on the dynamic changes and clonal amplification of the immune repertoire, requiring accurate differentiation between vaccine-induced clonal amplification and contaminating amplification introduced during the experiment.
[0183] The system received sequencing data from multiple volunteers at different time points (e.g., before and after vaccination). The model may calculate a higher risk of contamination between samples from the same volunteer at different time points, as these samples are typically processed consecutively on the same PCR plate. However, the overall batch noise baseline calculated based on all volunteer samples was relatively low, indicating that the overall quality control of this large-scale testing experiment was good. Historical contamination causality plots may reveal that certain operators or workstations have historically been associated with higher contamination outputs.
[0184] Low-frequency filtering within lanes: In this case, there may be no clear positive reference sample. The system can rely on predefined immune repertoire standards or use the same volunteer's pre-vaccination sample as a baseline reference to estimate the initial frequency ratio threshold. Combined with the calculated inter-sample contamination risk coefficient, the threshold is dynamically set to effectively filter "shadow" clones caused by cross-contamination within the same lane.
[0185] Inter-lane sample ratio filtering: The system detects a clone that is widely present in multiple samples from different volunteers, but its distribution is abnormally concentrated in certain lanes. Since the overall noise baseline of this batch is low, the system sets a relatively lenient but still meaningful fold-over threshold to identify and remove clones that may be abnormally spread in specific lanes due to aerosol contamination or reagent contamination. These clones do not conform to the biological pattern of vaccine-specific response.
[0186] Nucleotide sequence diversity verification: This step is also applied to remove non-functional sequences introduced by PCR or sequencing errors.
[0187] The system detected an abnormally high clonal overflow index in a sample, indicating that high-frequency clones in that sample had contaminated other samples in the same lane. Combined with its high overall contamination risk score, the system issued a "warning" for that sample, alerting researchers. This helps determine the extent to which the observed clonal amplification in the sample is genuine, or whether it is data distortion due to its strong contamination source characteristics, providing important quality control references for accurately assessing vaccine immunogenicity.
[0188] After comparing the data with the background contamination library, some public clone sequences that repeatedly appear in irrelevant research projects and are not related to the vaccine (e.g., some common clones related to baseline immune status) may be removed to ensure that the analysis focuses on vaccine-specific responses.
[0189] The novel contamination patterns, sample metadata, and corresponding contamination relationships identified in this large-scale clinical trial provide invaluable training data for the contamination risk prediction model, making its contamination predictions more accurate for future similar large-scale clinical research projects. Newly discovered background clones have been updated to the background database, enabling the system to achieve closed-loop learning and performance improvement.
[0190] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.
Claims
1. A method for filtering sequence contamination between samples in high-throughput sequencing of immune repertoires, characterized in that: include: S11: Preprocessing and contamination risk assessment, receiving sequencing data, sample metadata and reference sample data; Based on sample metadata, the pollution risk coefficients between samples and the overall data noise baseline of the batch are calculated through a pollution risk prediction model, and a historical pollution causal graph is constructed. S12: Dynamic core filtering, which uses reference sample data, contamination risk coefficient and data noise baseline to dynamically adjust the filtering threshold and standard, and performs three-stage sequence filtering on the sequencing data in sequence, including low-frequency filtering within lanes, sample ratio filtering between lanes and nucleotide sequence diversity verification. S13: Pollution diagnosis and sample-level early warning. Based on the filtering results, calculate the multi-dimensional filtering indicators for each sample, and generate sample-level pollution risk scores and graded early warnings through the sample pollution risk model. S14: Cross-batch collaborative filtering compares the filtered sequences with an updatable cross-project background contamination library and removes matching background contamination sequences; S15: System feedback and update. The new pollution patterns and background clone sequences identified in this filtering are fed back and updated to the pollution risk prediction model, historical pollution causal diagram and cross-project background pollution database.
2. The method for filtering sequence contamination between samples in high-throughput sequencing of immune repertoires according to claim 1, characterized in that: The pretreatment and pollution risk assessment steps specifically include: S21: Receive the input sequencing data, sample metadata, and sequencing data of a reference sample composed of known clones; S22: The pollution risk prediction model trained based on historical data calculates the pollution risk coefficient between any two samples according to the input sample metadata, and calculates the data noise baseline representing the overall background noise level of the current batch of experiments. S23: Based on historical batch data, using samples as nodes and the identified pollution directions as directed edges, construct a pollution causal graph that displays the pollution propagation path.
3. The method for filtering sequence contamination between samples in high-throughput sequencing of immune repertoires according to claim 2, characterized in that: In the pollution risk prediction model trained on historical data, when calculating the pollution risk coefficient between any two samples based on the input sample metadata, and calculating the data noise baseline representing the overall background noise level of the current batch of experiments, the specific steps include: S31: Feature normalization, which normalizes the numerical features in the input sample metadata, including sequencing depth, library concentration and PCR cycle number. S32: Individual pollution risk score calculation. The normalized features and the spatial location information of the sample in the lane are input into the pollution risk prediction model. The pollution risk prediction model calculates the individual pollution risk score for each sample through pre-trained weights. S33: Calculation of the contamination risk coefficient between samples. Based on the individual contamination risk scores of each pair of samples and their proximity relationship in the experimental operation, the contamination risk coefficient between any two samples is calculated. S34: Overall data noise baseline calculation. Based on the individual contamination risk scores of all samples in the current batch, a comprehensive index representing the overall background noise level of this batch is calculated, namely the overall data noise baseline.
4. The method for filtering sequence contamination between samples in high-throughput sequencing of immune repertoires according to claim 3, characterized in that: When constructing a pollution causal graph that displays the pollution propagation path based on historical batch data, using samples as nodes and identified pollution directions as directed edges, the specific steps include: S41: Data input, receiving historical batch data, including clonal composition information for each historical sample, as well as potential contamination pairs and their directions identified from historical analysis; S42: Graph structure definition, define a directed graph as a pollution causal graph; S43: Contamination Relationship Identification: For any two historical samples, determine whether there is a potential contamination relationship between the two historical samples; S44: Directed edge construction and weight assignment. For the identified contamination relationship, a directed edge is constructed and a weight is assigned to the constructed directed edge. S45: Pollution source identification and analysis. Based on the constructed pollution causal graph, each vertex in the graph is analyzed, and its degree as a pollution source and victim is evaluated by comparing the total weight of its outgoing and incoming edges.
5. The method for filtering sequence contamination between samples in high-throughput sequencing of immune repertoires according to claim 4, characterized in that: Low-frequency filtering within the swimming lane specifically involves: S51: For all samples within the same sequencing lane, merge all their clone sequences and calculate the frequency of each clone in each sample; S52: Analyze the data of the reference sample in the same lane and identify all known true clones; for each known clone, calculate the frequency ratio of its contamination to any other sample in the same lane, and dynamically determine the initial frequency ratio threshold based on the distribution of all calculated frequency ratio values for all known clones. S53: Based on the calculated pollution risk coefficient, the initial threshold is adjusted to obtain the personalized final frequency ratio threshold used for this comparison; S54: For any two samples within a lane and any clone they share, if their frequency ratio satisfies: ; Then, clone c in the low-frequency sample is determined to be caused by contamination from the high-frequency sample, and clone c is removed from the clone list of sample j. To determine the frequency of clone c in sample i, Let c be the frequency of clone c in sample j.
6. The method for filtering sequence contamination between samples in high-throughput sequencing of immune repertoires according to claim 5, characterized in that: The specific method for inter-lane sample ratio filtering is as follows: S61: For each clone sequence that still exists after filtering within the lane, calculate the proportion of its occurrence in each sequencing lane. S62: Based on the calculated data noise baseline representing the overall background noise level of the current batch, the exceedance multiple threshold for judging the proportion of abnormality is dynamically determined through a preset mapping function. S63: Contamination clone determination. For a specific clone and lane, determine whether it is an abnormally spread contaminant in that lane by comparing the relationship between its appearance ratio in the sample in that lane and the average appearance ratio in the sample in other lanes. S64: For a clone that is determined to be contaminated in the lane, the clone is sorted from low to high frequency among all samples in which the clone appears. In this order, the clone sequence is removed from the samples with the lowest frequency. After each removal, the sample occurrence ratio of the clone in the lane is recalculated until the sample occurrence ratio no longer meets the contamination determination criteria.
7. The method for filtering sequence contamination between samples in high-throughput sequencing of immune repertoires according to claim 6, characterized in that: Pollution diagnosis and sample-level early warning specifically include: S71: For each sample, based on the filtering results, calculate the sequence loss rate index, clone overflow index, and filtered clone characteristic index in the three stages of low-frequency filtering within lanes, sample ratio filtering between lanes, and nucleotide sequence diversity verification. S72: Input the calculated multi-dimensional indicators into the sample pollution risk model, and calculate the comprehensive pollution risk score for each sample through a preset weighted fusion algorithm; S73: Based on the calculated comprehensive pollution risk score of the samples, the samples are divided into different pollution risk levels, and corresponding early warning information is generated.
8. The method for filtering sequence contamination between samples in high-throughput sequencing of immune repertoires according to claim 7, characterized in that: Cross-batch collaborative filtering specifically includes: S81: Based on historical batch data, construct and continuously update a cross-project background contamination library. The background contamination library records low-frequency cloning sequences and their meta-information that repeatedly appear in negative controls and irrelevant samples. S82: Compare all the clone sequences retained after filtering in the current batch with the clone sequences recorded in the background contamination library; S83: If a clone sequence in the current batch matches a sequence in the background contamination library, the clone is determined to be a systematic background contamination and is removed from the data in the current batch.
9. The method for filtering sequence contamination between samples in high-throughput sequencing of immune repertoires according to claim 8, characterized in that: System feedback and updates specifically include: S91: Add the newly obtained sample-to-pollution relationship data and its corresponding sample metadata to the historical training dataset, and use the updated dataset to periodically retrain the pollution risk prediction model. S92: The identified contamination relationships between samples are added as new directed edges and their weights to the historical contamination causal graph; S93: After evaluating residual clones in samples marked as high risk, as well as clone sequences that are highly similar to but not completely matched with the current background contamination library, according to preset rules, the clones that meet the criteria are added as new background contamination sequences to the cross-project background contamination library.
10. A sequence contamination filtering system for high-throughput sequencing of immune repertoire samples, used to implement the sequence contamination filtering method for high-throughput sequencing of immune repertoire samples according to any one of claims 1-9, characterized in that it comprises: The data receiving and preprocessing module is used to receive sequencing data, sample metadata, and reference sample data. The pollution risk assessment module is used to calculate the pollution risk coefficient between samples and the overall batch data noise baseline based on sample metadata through a pollution risk prediction model, and to construct and update historical pollution causal graphs. The dynamic core filtering module is used to dynamically adjust the filtering threshold and standard using reference sample data, contamination risk coefficient and data noise baseline, and sequentially perform low-frequency filtering within lanes, sample ratio filtering between lanes and nucleotide sequence diversity verification on the sequencing data. The pollution diagnosis and early warning module is used to calculate multi-dimensional filtering indicators for each sample based on the filtering results, and generate sample-level pollution risk scores and graded early warnings through the sample pollution risk model. The cross-batch collaborative filtering module is used to compare the filtered sequences with the background contamination library and remove matching background contamination sequences. The feedback and update module is used to feed back and update the identified new pollution patterns and background clone sequences to the pollution risk prediction model, historical pollution causal diagram and background pollution database.
Citation Information
Patent Citations
Gene detection data cleaning method and system based on artificial intelligence
CN119418762A
Ecological restoration effect evaluation method for micro-polluted river
CN119721857A