Method, device, apparatus, medium and product for detecting contamination of sample

By determining the read count of Y chromosome probes in samples and using the negative binomial distribution maximum likelihood estimation technique, the accuracy of low-abundance contamination detection in existing technologies is insufficient, achieving refined quantification and improved stability of extremely low-level exogenous DNA.

CN121747701APending Publication Date: 2026-03-27SHANGHAI WEIHE MEDICAL LAB CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately detect sample contamination at low abundance levels, especially in monitoring exogenous DNA in female blood samples where uncertainty exists. Traditional methods lack effective modeling and targeted optimization for low-complexity regions, resulting in insufficient detection sensitivity and reliability.

Method used

The first read count of the Y chromosome probe in the first bodily fluid sample and the second read count in the second bodily fluid sample are determined by computing equipment. Using the negative binomial distribution maximum likelihood estimation technique, combined with statistical analysis of multiple groups of ideal male bodily fluid samples, a model for determining the level of contamination is established, so as to achieve refined quantification of extremely low levels of exogenous DNA and reduce the risk of misjudgment.

Benefits of technology

It significantly improves the sensitivity and accuracy of sample contamination detection, enabling precise quantification of exogenous DNA at extremely low contamination levels, reducing the risk of misjudgment, and improving data reliability and detection stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121747701A_ABST
    Figure CN121747701A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a method, a device, equipment, a medium and a product for detecting sample pollution. The method includes determining a first read count for a Y chromosome probe in a first body fluid sample. The method further includes determining a desired second read count of the Y-chromosome probe in a second bodily fluid sample. The method further includes determining a contamination level for the first bodily fluid sample based on the first read count and the second read count. Through the method, high-sensitivity pollution detection of trace male-derived DNA in the first body fluid sample can be realized under the condition of not increasing additional sequencing cost and experimental process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure generally relate to the field of contamination detection, and specifically to methods, apparatus, devices, media, and products for detecting contamination in samples. Background Technology

[0002] With the continuous maturation of high-throughput sequencing technology, genomic testing has been widely applied in clinical diagnosis, genetic disease screening, and early tumor assessment. The rapid development of multi-omics testing methods has made it possible to extract more genetic information from limited biological samples. The overall industry trend is towards higher sensitivity, broader targeting range, and more stable data quality to meet the dual demands of accuracy and information content in medical testing.

[0003] Driven by the concept of precision medicine, sequencing data quality control tools have become increasingly professional and automated, demonstrating high reliability, particularly in complex sample analysis and low-abundance signal identification. Current technological advancements focus on improving capture efficiency, optimizing sequence analysis models, and enhancing noise suppression capabilities, leading to continuous improvements in data consistency and clinical interpretability of test results. This provides a solid technological foundation for the further popularization of genome sequencing technology. Summary of the Invention

[0004] Embodiments of this disclosure provide a method, apparatus, device, medium, and product for detecting sample contamination.

[0005] According to a first aspect of this disclosure, a method for detecting sample contamination is provided. The method includes determining a first read count of a Y chromosome probe in a first bodily fluid sample. The method also includes determining a second read count of the Y chromosome probe in a second bodily fluid sample. The method further includes determining a contamination level for the first bodily fluid sample based on the first and second read counts.

[0006] According to a second aspect of this disclosure, an apparatus for detecting sample contamination is provided. The apparatus includes a first read count determination module configured to determine a first read count of a Y chromosome probe in a first bodily fluid sample; a second read count determination module configured to determine a second read count of the Y chromosome probe in a second bodily fluid sample; and a contamination level determination module configured to determine the contamination level of the first bodily fluid sample based on the first and second read counts.

[0007] In a third aspect of this disclosure, an electronic device is provided, including at least one processor; and a storage device for storing at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the method according to the first aspect of this disclosure.

[0008] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the method according to a first aspect of this disclosure.

[0009] In a fifth aspect of this disclosure, a computer program product is provided. This computer program product includes a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.

[0010] It should be understood that the content described in this section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] The above and other objects, features and advantages of this disclosure will become more apparent from the accompanying drawings, in which like reference numerals generally denote like parts.

[0012] Figure 1 The illustration shows a schematic diagram of an example environment in which some embodiments of the present disclosure may be implemented;

[0013] Figure 2 The illustration shows a schematic diagram of an example method for detecting sample contamination according to some embodiments of the present disclosure;

[0014] Figure 3 The illustration shows a flowchart of an example process for probe design and verification according to some embodiments of the present disclosure;

[0015] Figure 4 The illustration shows a flowchart of an example process for determining pollution levels according to some embodiments of the present disclosure.

[0016] Figure 5 The illustration shows a schematic diagram of an example of Y chromosome probe coverage according to some embodiments of the present disclosure;

[0017] Figure 6 The illustration shows a schematic diagram of an example of a pollution assessment level according to some embodiments of the present disclosure;

[0018] Figure 7 The illustration shows a schematic block diagram of an apparatus for detecting sample contamination according to some embodiments of the present disclosure;

[0019] Figure 8 A schematic block diagram of an example device suitable for implementing various embodiments of the present disclosure is illustrated. Detailed Implementation

[0020] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0021] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0022] For example, upon receiving a user's proactive request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0023] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0024] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0025] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0026] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0027] Current methods for detecting contamination in blood samples primarily rely on inferences based on allele frequency changes at single nucleotide polymorphism (SNP) sites. These methods largely stem from experience with cross-contamination analysis of tumor tissue samples. However, when applied to scenarios with low DNA content, short fragments, and no paired samples to support the contamination, the contamination signal is easily masked by background noise. Furthermore, existing research largely focuses on whole-genome or exome sequencing data, and the robustness and sensitivity of their algorithms are not yet fully adapted to methylation capture sequencing data, significantly limiting reliable detection capabilities at extremely low contamination levels.

[0028] On the other hand, while relying on Y chromosome signals for sample contamination tracking has significant advantages in certain scenarios, the Y chromosome's biological characteristics, such as numerous repetitive sequences and low sequence complexity, coupled with the further loss of base information caused by bisulfite treatment, mean that conventional probe design strategies often suffer from insufficient coverage and poor specificity in these regions, leading to unstable experimental capture efficiency. Traditional methods lack effective modeling and targeted optimization for low-complexity regions, making it difficult to support high-sensitivity detection requirements, especially exhibiting strong uncertainty in monitoring extremely low levels of contamination in female blood samples.

[0029] Therefore, embodiments of this disclosure propose a method for detecting sample contamination. In this method, a computing device determines a first read count of the Y chromosome probe in a first bodily fluid sample. The computing device then determines a second read count of the Y chromosome probe in a second bodily fluid sample. The computing device further determines the contamination level of the first bodily fluid sample based on the first and second read counts. This method significantly improves the sensitivity of contamination detection while ensuring detection accuracy, enabling precise quantitative tracing of extremely low levels of exogenous DNA, effectively reducing the risk of false positives, and improving overall detection level and data reliability.

[0030] The embodiments of this disclosure will now be described in further detail with reference to the accompanying drawings. Figure 1 An example environment in which the apparatus and / or methods of embodiments of the present disclosure may be implemented is illustrated. In environment 100, computing device 102 can generate a contamination determination result corresponding to the true contamination level of the sample based on the Y chromosome probe read count in a first bodily fluid sample and the expected read count of these Y chromosome probes constructed in a second bodily fluid sample. In some embodiments, the first bodily fluid sample may be a female bodily fluid sample, and the second bodily fluid sample may be a male bodily fluid sample.

[0031] Examples of computing device 102 include, but are not limited to, any suitable gene testing device, personal computer, server computer, handheld or laptop device, mobile device, multiprocessor system, consumer electronics, minicomputer, mainframe computer, distributed computing environment including any of the above systems or devices, etc.

[0032] like Figure 1 As shown, the computing device 102 can be used to perform sample contamination identification and quantitative analysis tasks based on Y chromosome probes. For the input bodily fluid sequencing data, the computing device can process different probe signals to construct a potential exogenous DNA distribution model in the first bodily fluid sample. Specifically, the first read count 104 represents the support reads of the Y chromosome probes actually measured in the first bodily fluid sample to be tested, reflecting the true manifestation of contamination signals in the sample; the second read count 106 is established based on the second bodily fluid sample, representing the signal response level of each probe in an ideal second sample, such as the signal response level in a pure male sample, and is an important theoretical basis for contamination inference. Through the fusion calculation of these two types of read counts, the computing device 102 can output a contamination level 108 to indicate whether and to what extent male-derived DNA is mixed in the sample to be tested.

[0033] To improve the accuracy of contamination identification, the computing device 102 can perform pre-quality control operations on the first read count 104, including cleaning and standardizing probe data with insufficient sequencing depth, excessive background noise, or abnormal alignment results. Simultaneously, when establishing the second read count 106, the computing device can statistically analyze the coverage of each Y chromosome probe based on measured data from multiple sets of ideal male bodily fluid samples to determine the corresponding expected coverage reading. This expected value serves as a priori parameter for the model, characterizing the normal signal intensity of the probe in male bodily fluid samples. By employing a multi-sample statistical approach, biases caused by individual differences, sequencing depth fluctuations, or changes in experimental conditions can be effectively reduced, resulting in higher stability, representativeness, and reproducibility. This ensures that the second read count 106 accurately reflects the true signal level of male-derived DNA, providing a reliable benchmark for probabilistic inference of contamination levels. Through refined acquisition and consistency calibration of dual-source parameters, the scientific validity and statistical robustness of the input data for contamination determination can be effectively ensured.

[0034] During contamination analysis, computing device 102 models the parametric relationship between the first read count 104 and the second read count 106 to obtain the contamination level 108. This modeling process can employ negative binomial maximum likelihood estimation to represent the randomness and discreteness of the probe signal under different contamination levels. This not only identifies obvious contamination but also maintains detection sensitivity even with extremely low exogenous DNA incorporation, enabling the computing device to accurately capture rare contamination signals against a certain sequencing read noise background and provide confident contamination level inferences.

[0035] In the entire analysis mechanism, the first read count 104 provides the real contaminated sample signal, and the second read count 106 provides the comparative information. The joint expression of the two in the model ensures the accuracy of the contamination determination result.

[0036] This method enables computing devices to output the final contamination level without the need for paired normal reference samples, achieving automated, real-time, and modular data quality control capabilities.

[0037] The above combination Figure 1 The following is a schematic diagram illustrating an example environment in which some embodiments of this disclosure may be implemented, in conjunction with... Figure 2 A schematic diagram illustrating an example method for detecting sample contamination according to some embodiments of the present disclosure. Figure 2 The method in can be derived from Figure 1 The computing device 102 or any suitable device in the system can be used for execution.

[0038] like Figure 2 As shown, in example method 200, at box 202, computing device 102 determines a first read count for a Y chromosome probe in a first bodily fluid sample. This read count represents sequencing reads associated with male-specific genetic markers detected in the first bodily fluid sample and can be used to reflect the presence of traces of exogenous DNA from males in the sample, enabling a preliminary assessment of the risk of contamination.

[0039] In some embodiments, the computing device 102 may preprocess the sequencing data during probe read acquisition, including removing low-quality alignments, anomalous mapping positions, and potential polymerase chain reaction (PCR) artifacts. The first read count obtained through this process more accurately reflects the true contribution of contamination, thereby reducing the risk of misjudgment due to noise.

[0040] In box 204, computing device 102 determines the second read count of the Y chromosome probe in a second bodily fluid sample. For example, the desired second read count is constructed based on bodily fluid samples from multiple men as a priori parameter of the model, describing the normal response level of each probe in normal male bodily fluid.

[0041] In some embodiments, the second read count can be generated by averaging, variance fitting, or distribution estimation of probe readings from multiple male control samples to obtain a more stable and representative parameter. The computing device 102 can also perform bias correction based on sequencing platform, library preparation reagents, and batch differences to ensure that the determined second read count is consistent across batches and laboratories.

[0042] At box 206, computing device 102 determines the contamination level for the first bodily fluid sample based on the first and second reading counts. The computing device treats the first reading count as a random variable following a negative binomial distribution, and its expected value is obtained by multiplying the contamination level by the second reading count. Based on this, the maximum likelihood estimation method is used to determine the contamination level that most likely results in the observed signal, and this level is output as the sample contamination assessment result.

[0043] In some embodiments, the computing device may use a first read segment count and a second read segment count to determine a maximum likelihood estimate of the contamination level for a first bodily fluid sample. For example, the maximum likelihood estimate may be obtained by maximizing the log-likelihood of a likelihood function formed based on the read segment counts. The computing device then uses the maximum likelihood estimate as the contamination level.

[0044] In some embodiments, multiple Y-chromosome probes, including those described above, are used when calculating the maximum likelihood estimate. Therefore, when determining the maximum likelihood estimate of the contamination level for the first bodily fluid sample, multiple first read counts for the multiple Y-chromosome probes in the first bodily fluid sample can be determined, wherein the multiple first read counts include first read counts. Additionally, a desired multiple second read counts for the multiple Y-chromosome probes in the second bodily fluid sample can be determined. The computing device can then use the multiple first read counts and the multiple second read counts to determine the maximum likelihood estimate of the contamination level for the first bodily fluid sample.

[0045] To enhance the reliability of contamination determination, the computing device 102 can also simultaneously output confidence indicators of the contamination results, such as likelihood function convergence, variance range, or confidence interval, enabling the testing to provide more interpretable and clinically relevant quality control results. Furthermore, the computing device can dynamically set risk levels based on the numerical distribution of contamination levels, used to identify high-risk samples and trigger additional testing or verification processes.

[0046] In some embodiments, the computing device 102 can perform batch analysis of target samples and make unified judgments based on statistical information of the second read segment counts to support large-scale sample contamination screening scenarios.

[0047] Before determining the contamination level, the computing device 102 can also perform a Y chromosome probe screening process to obtain a set of Y chromosome probes for subsequent detection. First, the computing device determines multiple candidate regions based on the Y chromosome sequence. The computing device 102 can evaluate the overlap ratio between the candidate regions and predetermined repetitive sequence regions. If the overlap ratio of a candidate region with a repetitive sequence region is higher than a threshold ratio, the candidate region is removed because it is not a special or unique region. Only regions with an overlap ratio lower than the threshold and clear sequence characteristics are identified as valid targets for probe design. Therefore, candidate regions with an overlap ratio lower than the threshold ratio can be identified as regions within a set of regions used to generate candidate Y chromosome features.

[0048] Then, several candidate probe sequences can be generated in this set of regions to form an initial candidate set. Each candidate probe is used to cover a potential male-specific genetic region, thereby ensuring sensitive detection capability of exogenous male DNA.

[0049] During the candidate probe generation stage, the computing device 102 can convert the sequences for regions within a set of regions into fully methylated and fully demethylated sequences, respectively, and truncate them according to a predetermined length and a fixed step size to form multiple probe sequences. At this point, fully methylated and fully demethylated sequences can be generated for each region within a set of regions. For example, the predetermined length can be set to 120 bp, and the fixed step size can be set to 10 bp to ensure more continuous coverage and compatibility with the experimental characteristics of methylation capture sequencing.

[0050] Subsequently, the computing device 102 determines multiple evaluation parameters for each candidate Y chromosome probe to rank them. In some embodiments, these evaluation parameters may include one or more of the following: thermodynamic parameters of the candidate Y chromosome probe, off-target risk assessment results of the candidate Y chromosome probe or off-target risk assessment results caused by potential matching with other regions of the genome, secondary structure thermodynamic parameters, and interaction assessment results of the candidate Y chromosome probe. Therefore, these evaluation parameters can be determined by determining one or more of the above parameters. By combining multi-dimensional parameter evaluation, the computing device can predict the signal performance and stability of the probe under real experimental conditions.

[0051] After obtaining the probe evaluation metrics, the computing device 102 assigns scores to all candidate Y chromosome probes based on these parameters and sorts them according to the scores to form a probe candidate pool. High-scoring probes indicate better performance in terms of experimental reliability, specificity, and signal consistency, and can be prioritized for contamination monitoring. This sorting mechanism effectively filters out probe sequences with poor performance or high potential risks, thereby improving the success rate and stability of subsequent detection processes. Alternatively, scoring can also be based on other information, such as melting temperature.

[0052] The computing device 102 then further filters the candidate probe pool based on the actual performance in the sample sequencing data to ensure the probes are operable in practical applications. For example, if a candidate Y chromosome probe has no significant reading in male control samples, it indicates insufficient capture efficiency and the candidate Y chromosome probe needs to be removed; if a probe has a significant reading in female control samples, it indicates potential background noise. For example, if the reading of a candidate Y chromosome probe in a female bodily fluid sample is greater than a threshold, the candidate Y chromosome probe is removed, such as removing candidate Y chromosome probes with readings greater than 5; if the coverage difference of a candidate Y chromosome probe between different male bodily fluid samples exceeds a threshold difference, it indicates inaccurate testing and the probe is removed. Therefore, these unsuitable candidate Y chromosome probes will be eliminated to reduce noise propagation.

[0053] This multi-stage screening process yields a set of highly specific and stable target Y chromosome probes. These probes not only exhibit a significant signal response to male-derived DNA but also show extremely low background expression in female samples, enabling them to provide high-confidence support for contamination monitoring.

[0054] This method demonstrates higher sensitivity, stronger robustness, and better engineering adaptability in methylation capture sequencing scenarios, which can significantly improve overall data reliability and clinical decision-making effectiveness.

[0055] The above combination Figure 2 Schematic diagrams illustrating example methods for detecting sample contamination, representing some embodiments of this disclosure, are described below. Figure 3 A flowchart describing an example process for probe design and verification according to some embodiments of this disclosure. Figure 3 Example method 300 in the example can be derived from Figure 1 The computing device 102 or any suitable device in the system can be used for processing.

[0056] like Figure 3As shown in Example Flow 300, the probe design and validation process for the Y chromosome region is illustrated. The computing device progressively performs sequence screening, probe attribute evaluation, and risk filtering based on the target sequence characteristics, thereby generating a set of candidate probes for experimental detection.

[0057] At box 302, the computing device acquires and parses the Y chromosome region sequence, initiating the Y chromosome probe design process. Then, at box 304, repetitive regions in the Y chromosome sequence are screened using genomic database annotation information. For example, if the overlap between a candidate sequence region and a repetitive sequence annotation region exceeds 20%, the region is directly removed to ensure the specificity and uniqueness of subsequent probe design. Next, the retained selected sequence regions can undergo target sequence methylation transformation. For example, for each target sequence, two types of methylation transformation sequences are constructed based on the methylation status of cytosine-phosphate-guanine (CpG) sites. One is a fully methylated sequence, for example, retaining all C at all CpG sites unchanged, while converting C at other non-CpG (CHH) sites to T. The other is a fully demethylated sequence, for example, converting all C sites to T. This step aims to simulate DNA sequences under different methylation states to aid in subsequent probe design.

[0058] Probe sequences can be designed for these fully methylated and fully demethylated sequences. For example, a tiling strategy can be used to design 120 bp probes from the methylated target sequence at a fixed step size of 10 bp. The probes cover both the positive (OT) and negative (OB) strands, thus each sequence position corresponds to four probes, representing fully methylated / fully demethylated and combinations of positive and negative strands, ensuring comprehensive and flexible probe coverage. These probes serve as candidate Y chromosome probes. For these candidate probes, the computational device performs thermodynamic parameter calculations at frame 306, including melting temperature (Tm) and free energy change (ΔG). The binding stability of the probes to the target sequence is quantitatively verified by simulating molecular binding strength under hybridization conditions. This step helps eliminate probes with insufficient binding ability or excessive binding that affects dynamic unwinding, ensuring accurate targeting in subsequent capture steps.

[0059] In addition, the computing device performs off-target risk assessment in box 308. At this point, a pre-defined tool can be used to align the methylated human genome, and alignment results with a bit-score > 90 are defined as high-risk off-target sites. The number of high-risk hits is counted for each probe, and probes with more than 50 high-risk hits are removed to reduce the possibility of non-specific binding. The number of hits refers to the number of candidate sequences / sites that meet thresholds such as similarity and matching length with the target sequence (e.g., gene, CpG island) during sequence alignment / database retrieval. In box 310, a secondary structure risk assessment is performed on the probes. At this point, the thermodynamic parameters of the secondary structure of candidate probes can be calculated, such as the melting temperature for hairpin structures, to screen out probe sequences that may form their own hairpin structures or other unfavorable secondary structures, ensuring the effectiveness of the probes under experimental conditions.

[0060] In box 312, probe interaction evaluation can be performed to obtain the interaction evaluation results. At this time, the probe sequences are compared pairwise to evaluate whether there is a tendency for complementary pairing between the probes. Probes that may form dimers or other interaction structures are eliminated to avoid mutual interference between probes affecting the detection effect.

[0061] Subsequently, in box 314, the candidate probe set was experimentally verified to confirm its repeatable capture performance under real sample conditions, enabling the probe set to reliably adapt to subsequent detection processes.

[0062] This method enables probe design and theoretical parameter detection, allowing the final Y chromosome probe to maintain significant coverage and good specificity in low-complexity regions.

[0063] The above combination Figure 3 A flowchart illustrating an example process for probe design and verification according to some embodiments of this disclosure is described below; in conjunction with... Figure 4 A flowchart describing an example process for determining pollution levels according to some embodiments of this disclosure. Figure 4 Example process 400 can be generated by Figure 1 The computing device 102 shown or any suitable device may be used for execution.

[0064] like Figure 4 As shown, Example Flow 400 constructs a contamination monitoring mechanism based on probe signal intensity. This mechanism is based on a validated set of Y chromosome probes and achieves continuous calculation of contamination levels by correlating male-specific signals with weak expression in mixed samples in a quasi-quantitative relationship.

[0065] In box 402, the computing device first performs experimental performance verification on the probe to confirm that it maintains stable capture efficiency and signal consistency under actual library construction and sequencing conditions, thereby ensuring the quality of model input. By comparing the background signal differences between ideal male and female samples, probes that are highly noise-sensitive or have low reproducibility can be excluded in advance, avoiding systematic bias in subsequent contamination estimation.

[0066] At box 404, the computing device removes probes with abnormal signals based on experimental coverage. For example, it removes crossover Y chromosome probes that have no read coverage; it removes probes with more than 5 reads coverage in any female sample; and it removes probes with significant differences in coverage among male samples.

[0067] In box 406, the computing device determines the relationship between probe coverage and sequencing volume in male samples. For example, the desired probe coverage parameter can be established based on multiple sets of male standard samples. , which represents the average detection level of each probe under ideal undiluted conditions (pure male samples).

[0068] Additionally, assuming X_i represents the number of supporting molecules (reads) for the i-th probe in a blood sample, it can be modeled using a negative binomial distribution: ,in For the dispersion parameter, probe Expected coverage , Let be the expected number of reads covered by probe i in an ideal male sample. This model can accurately characterize the excessive dispersion characteristics under low-coverage sequencing noise. In box 408, construct a pollution monitoring model. The probability mass function of a negative binomial distribution can be constructed as follows: Under the assumption that each probe is independent, Substituting the values, we can obtain the population likelihood function. The pollution level estimate is then obtained by log-likelihood maximization. By aggregating consensus among multiple probes to suppress fluctuations in a single probe, the sensitivity and robustness of contamination detection in the context of extremely low male DNA levels are effectively improved.

[0069] When new sample data is input at box 410, the computing device maps its sequencing support to the established model and outputs the contamination level determination result at box 412. This determination not only includes the contamination ratio but also quantifies the reliability of the result based on the distribution convergence, enabling the computing device to upgrade the contamination determination from a label-based approach to a continuous risk assessment capability.

[0070] The above combination Figure 4A flowchart illustrating an example process for determining pollution levels according to some embodiments of this disclosure is described below. Figure 5 A schematic diagram illustrating an example of Y chromosome probe coverage according to some embodiments of the present disclosure.

[0071] For example, strictly quality-controlled, uncontaminated cell-free DNA (cfDNA) samples can be selected as baseline controls, and sample contamination scenarios can be simulated through artificial mixing. Specifically, male cfDNA samples are mixed into female cfDNA samples at proportions of 1%, 0.5%, 0.1%, 0.05%, and 0.01%, respectively, thus constructing a series of simulated datasets with different contamination levels. Since female samples naturally do not contain the Y chromosome, ideally, the coverage of the Y chromosome probe in their sequencing results should be zero.

[0072] Based on this, the Y chromosome coverage of each mixed sample was systematically examined. The results showed that: 1) in pure female cfDNA samples, the Y chromosome probe had a certain noise signal; 2) in mixed samples, the number of Y chromosome probe reads was significantly positively correlated with the incorporation ratio, that is, as the proportion of male cfDNA increased, the detection signal of the Y chromosome also increased linearly.

[0073] The experimental results clearly validate that Y chromosome coverage can serve as a sensitive and specific indicator for contamination detection. Even at a mixing level as low as 0.01%, Y chromosome reads can still be observed, providing a methodological basis for contamination monitoring of actual cfDNA samples. Figure 5 As shown in Example 500, the coverage distribution of the Y chromosome-specific probe in samples with different contamination gradients is illustrated. This figure visually demonstrates the signal presentation characteristics of the probe under different male DNA incorporation ratios, validating the signal-to-noise separation and sensitive response performance of the Y chromosome probe in contamination detection scenarios. With changes in incorporation ratio, the coverage color gradient exhibits significant differences, allowing direct observation of the impact of contamination levels on the probe detection signal.

[0074] When the sample source is pure female cell-free DNA, the probe signals in the figure are all at extremely low background levels, showing a discrete light-colored distribution, indicating that the probe set has good sex specificity and an extremely low false positive rate. At the same time, as the proportion of male DNA incorporation gradually increases, the coverage in the corresponding probe region is significantly enhanced, forming a continuous signal band from light to dark, indicating that the accumulation of contamination contribution shows a consistent trend with the actual degree of contamination.

[0075] Under low-contamination conditions (such as 0.05% or even lower), a stable signal exceeding the background noise was observed in the probe region, demonstrating that the probe detection scheme can still reliably identify contamination scenarios at the tens of thousands level. The signal contrast caused by concentration differences was clearly mapped onto the heatmap, enabling clear distinction between different contamination gradients and showcasing a highly sensitive quantitative detection capability.

[0076] The coverage in high-mixture samples rapidly increased into the darker regions, reflecting a strong positive correlation between probe capture ability and actual male DNA content. This correlation not only validates the effectiveness of the probe design strategy but also supports a reliable inference of the potential contamination ratio from the coverage level.

[0077] The visualization of signal differences in this figure demonstrates that the Y chromosome probe set can accurately indicate the source and extent of contamination in a complex cellular DNA background, effectively improving sample quality control capabilities.

[0078] In some embodiments, probes can be selected and filtered. For example, when systematically screening probes for the Y chromosome region, multiple filtering criteria can be used to ensure that only probes with stable detection performance and high specificity are retained for subsequent analysis. In this process, continuous probe merging can be performed: adjacent probes on the Y chromosome are merged to reduce redundancy and improve signal consistency. For example, this step yields 151 candidate probes. Then, coverage screening is performed. At this point, effective sequencing read coverage is required in all male samples, ultimately resulting in 121 probes with coverage. Female background filtering can also be performed: considering that female samples should not contain Y chromosome signals, probes with read coverage greater than 5 in any female sample are removed to reduce false positive signals introduced by non-specific capture or noise, leaving 111 probes. Additionally, probe validity checks can be performed: probes with missing values ​​(NA value) in any male sample are removed to ensure that the retained probes have stable detection signals in male samples, ultimately retaining 104 probes. Capture performance optimization: probes with unstable capture efficiency or abnormal performance are eliminated, resulting in 99 high-quality Y chromosome probes. Through the above-described stratified filtering process, a set of 99 highly specific and stable Y chromosome probes was successfully obtained, which can be used for subsequent cfDNA sample contamination detection. See Table 1 below. Table 1. Probe screening steps and number of probes to be retained

[0079] The following is combined Figure 6 A schematic diagram illustrating examples of pollution assessment levels according to some embodiments of this disclosure.

[0080] like Figure 6 As shown, Example 600 illustrates the correlation between the true contamination level and the contamination level estimated based on the Y chromosome probe under different sample mixing conditions. Each point represents the detection results of DNA from different male sources mixed into cell-free DNA in female cells. Different sample batches are distinguished by color intensity to verify the stability of the model across sample environments.

[0081] exist Figure 6 The figure shows the relationship between the actual contamination percentage (x-axis) and the corresponding assessed contamination level (y-axis). All data points for all samples are distributed near the reference diagonal, indicating a high degree of consistency between the contamination estimation results and the actual mixing ratio. The assessed contamination level increases significantly as the contamination percentage increases from 0.01% to 1%. Different contamination levels are distinguishable from each other; contamination at the ten-thousandth percentile level (0.05%) can be detected, demonstrating the method's extremely high sensitivity. Specifically, within 10... -3 Up to 10 -2 The data points are closely clustered on both sides of the diagonal, indicating that the model has good quantitative accuracy at common pollution levels and can reliably reflect the actual pollution intensity of the samples. In the ultra-low pollution range (10... -4 Even at magnitudes below zero, multiple data points remain close to the diagonal trend line, further demonstrating the model's ability to identify extremely low pollution levels. The fact that the bias remains within a reasonable range indicates that the negative binomial maximum likelihood estimation method effectively distinguishes background noise from actual pollution signals.

[0082] This method can maintain excellent pollution level estimation accuracy even under pollution levels of 10,000, enabling reliable quantification of pollution levels and effectively compensating for the performance deficiencies of existing methods in identifying low-abundance pollution.

[0083] Figure 7 The illustration shows a schematic block diagram of an apparatus for detecting sample contamination according to some embodiments of the present disclosure. Figure 7 As shown, device 700 can Figure 1 The device 700 is implemented in a computing device 102 and includes a first read count determination module 702 configured to determine a first read count of the Y chromosome probe in a first body fluid sample; a second read count determination module 704 configured to determine a second read count of the Y chromosome probe in a second body fluid sample; and a contamination level determination module 706 configured to determine the contamination level of the first body fluid sample based on the first read count and the second read count.

[0084] In some embodiments, the contamination level determination module 706 includes: a contamination level maximum likelihood estimation module configured to determine a maximum likelihood estimate of the contamination level for a first body fluid sample based on a first read segment count and a second read segment count; and a contamination level determination module configured to use the maximum likelihood estimate as the contamination level.

[0085] In some embodiments, where the Y chromosome probe is one of a plurality of Y chromosome probes, the maximum likelihood estimation module for contamination level includes: a plurality of first read count determination modules configured to determine a plurality of first read counts for the plurality of Y chromosome probes in a first body fluid sample; a plurality of second read count determination modules configured to determine a plurality of expected second read counts for the plurality of Y chromosome probes in a second body fluid sample; and a maximum likelihood estimation determination module configured to determine a maximum likelihood estimate of the contamination level for the first body fluid sample based on the plurality of first read counts and the plurality of second read counts.

[0086] In some embodiments, the apparatus 700 further includes: a candidate probe determination module configured to determine a plurality of candidate Y chromosome probes for a set of regions of a Y chromosome sequence; an evaluation parameter determination module configured to determine a plurality of evaluation parameters for candidate Y chromosome probes among the plurality of candidate Y chromosome probes; a probe candidate pool generation module configured to sort the plurality of candidate Y chromosome probes based on the plurality of evaluation parameters to form a probe candidate pool; and a chromosome probe filtering module configured to filter the plurality of candidate Y chromosome probes in the probe candidate pool to determine the Y chromosome probe.

[0087] In some embodiments, the chromosome probe filtering module uses at least one of the following to filter multiple candidate Y chromosome probes in the probe candidate pool: a second sample probe filtering module configured to remove candidate Y chromosome probes that have no readings in a second body fluid sample; a first sample probe filtering module configured to remove candidate Y chromosome probes whose readings in a first body fluid sample are greater than a threshold; and a coverage difference probe filtering module configured to remove candidate Y chromosome probes whose coverage differences in different second body fluid samples are greater than a threshold difference.

[0088] In some embodiments, the apparatus 700 further includes: an overlap ratio determination module configured to determine the overlap ratio between a candidate region and a predetermined repetitive sequence region among a plurality of candidate regions of the Y chromosome sequence; and an overlap ratio screening module configured to determine the candidate region as a region of a set of regions in response to an overlap ratio being less than a threshold ratio.

[0089] In some embodiments, the apparatus 700 further includes a candidate region removal module configured to remove candidate regions in response to an overlap ratio greater than or equal to a threshold ratio.

[0090] In some embodiments, the candidate probe determination module includes: a sequence conversion module configured to convert sequences targeting regions within a set of regions into fully methylated sequences and fully demethylated sequences; and a candidate probe generation module configured to determine candidate Y chromosome probes targeting fully methylated sequences and fully demethylated sequences based on a predetermined length and a fixed step size.

[0091] In some embodiments, the predetermined length is 120 bp and the fixed step size is 10 bp.

[0092] In some embodiments, the evaluation parameter determination module includes at least one of the following: a thermodynamic parameter determination module configured to determine thermodynamic parameters for the candidate Y chromosome probe; an off-target risk assessment result determination module configured to determine an off-target risk assessment result for the candidate Y chromosome probe; a secondary structure thermodynamic parameter determination module configured to determine secondary structure thermodynamic parameters for the candidate Y chromosome probe; and an interaction assessment result determination module configured to determine an interaction assessment result for the candidate Y chromosome probe.

[0093] In some embodiments, the probe candidate pool generation module includes: a chromosome probe score determination module configured to determine multiple scores for multiple candidate Y chromosome probes based on multiple evaluation parameters; and a chromosome probe sorting module configured to sort the multiple candidate Y chromosome probes based on the multiple scores.

[0094] Figure 8 A schematic block diagram of an example device 800 that can be used to implement embodiments of the present disclosure is shown. Figure 1 The computing device 102 can be implemented using device 800. As shown, device 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 802 or loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 can also store various programs and data required for the operation of device 800. CPU 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 807 is also connected to bus 804.

[0095] Multiple components in device 800 are connected to I / O interface 807, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0096] The various processes and handling described above, such as method 200, can be executed by processing unit 801. For example, in some embodiments, method 200 can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by CPU 801, one or more actions of the example method 200 described above can be performed.

[0097] This disclosure can be a method, apparatus, system, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.

[0098] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0099] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0100] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0101] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0102] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0103] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0104] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0105] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for detecting sample contamination, comprising: Determine the count of the first read segment of the probe against the Y chromosome in the first bodily fluid sample; Determine the second read count of the Y chromosome probe in the second body fluid sample; as well as Based on the first read segment count and the second read segment count, the contamination level of the first bodily fluid sample is determined.

2. The method of claim 1, wherein determining the contamination level of the first bodily fluid sample comprises: Based on the first and second read counts, determine the maximum likelihood estimate of the contamination level for the first bodily fluid sample; as well as The maximum likelihood estimate is taken as the pollution level.

3. The method of claim 1, wherein the Y chromosome probe is one of a plurality of Y chromosome probes, and determining the maximum likelihood estimate of the contamination level for the first bodily fluid sample based on the first read count and the second read count comprises: Determine the count of multiple first reads against multiple Y chromosome probes in the first bodily fluid sample; Determine the desired number of second reads of the plurality of Y chromosome probes in a second body fluid sample; as well as Based on the plurality of first read segment counts and the plurality of second read segment counts, a maximum likelihood estimate of the contamination level for the first bodily fluid sample is determined.

4. The method according to claim 1, further comprising: Identify multiple candidate Y chromosome probes targeting a set of regions of the Y chromosome sequence; For each candidate Y chromosome probe among the plurality of candidate Y chromosome probes, a plurality of evaluation parameters are determined for the candidate Y chromosome probe; as well as Based on the multiple evaluation parameters, the multiple candidate Y chromosome probes are sorted to form a probe candidate pool; The candidate Y chromosome probes in the probe candidate pool are filtered to determine the Y chromosome probe.

5. The method of claim 4, wherein filtering the plurality of candidate Y chromosome probes in the probe candidate pool comprises: The plurality of candidate Y chromosome probes in the probe candidate pool are filtered using at least one of the following: Remove candidate Y chromosome probes that did not have a reading in the second body fluid sample; Remove candidate Y chromosome probes whose readings are greater than the threshold in the first bodily fluid sample; or Candidate Y chromosome probes whose coverage differences in different second body fluid samples exceed a threshold difference are removed.

6. The method of claim 4, further comprising: Determine the overlap ratio between candidate regions and predetermined repetitive sequence regions among multiple candidate regions of the Y chromosome sequence; as well as In response to the overlap ratio being less than a threshold ratio, the candidate region is determined as a region within the set of regions.

7. The method according to claim 6, further comprising: In response to the overlap ratio being greater than or equal to the threshold ratio, the candidate region is removed.

8. The method of claim 4, wherein determining a plurality of candidate Y chromosome probes for a set of regions of the Y chromosome comprises: The sequences targeting regions within the set of regions are converted into fully methylated and fully demethylated sequences; as well as Based on a predetermined length and a fixed step size, the plurality of candidate Y chromosome probes for the fully methylated sequence and the fully demethylated sequence are determined.

9. The method of claim 8, wherein the predetermined length is 120 bp and the fixed step size is 10 bp.

10. The method of claim 4, wherein determining the plurality of evaluation parameters for the candidate Y chromosome probe includes at least one of the following: Determine the thermodynamic parameters for the candidate Y chromosome probe; Determine the off-target risk assessment results for the candidate Y chromosome probe; Determine the secondary structure thermodynamic parameters for the candidate Y chromosome probe; or Determine the interaction assessment results for the candidate Y chromosome probes.

11. The method of claim 4, wherein ranking the plurality of candidate Y chromosome probes to form a probe candidate pool based on the plurality of evaluation parameters comprises: Based on the multiple evaluation parameters, multiple scores are determined for the multiple candidate Y chromosome probes; as well as Based on the scores, the candidate Y chromosome probes are ranked.

12. An apparatus for detecting sample contamination, comprising: The first read segment count determination module is configured to determine the first read segment count against the Y chromosome probe in the first body fluid sample. The second read segment counting determination module is configured to determine the second read segment count of the Y chromosome probe in a second body fluid sample; as well as The contamination level determination module is configured to determine the contamination level for the first bodily fluid sample based on the first read segment count and the second read segment count.

13. An electronic device, comprising: At least one processor; as well as A memory for storing at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the method according to any one of claims 1-11.

14. A computer-readable storage medium having a computer program stored thereon, the computer program implementing the method according to any one of claims 1-11 when executed by a processor.

15. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-11.