Array-based targeted copy number detection

Cross-sample calibration using a reference signal distribution improves the precision of detecting target genomic regions in samples with contamination or uneven concentration by addressing signal interference and noise.

HK40135227APending Publication Date: 2026-07-17ILLUMINA INC

Patent Information

Authority / Receiving Office
HK · HK
Patent Type
Applications
Current Assignee / Owner
ILLUMINA INC
Filing Date
2026-06-11
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing methods for detecting target genomic regions in samples with contamination or uneven concentration face challenges in accurately determining variations due to signal interference and noise.

Method used

A method involving cross-sample calibration using a reference sample to construct a reference signal distribution, followed by calibrating intensity signals from input samples, and generating aggregated calibration signals to detect variations in target genomic regions.

Benefits of technology

Enhances the accuracy of detecting target genomic region variations by mitigating signal interference and noise, thereby improving detection precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Array-based targeted copy number detection, for instance detection on contaminated and / or variable concentration samples, includes obtaining a collection of intensity signals from assays of a set of input samples, performing a cross-sample calibration on the intensity signals based on reference sample(s), which calibration includes constructing a reference signal distribution based on intensity signals of the reference sample(s) and for one or more input samples calibrating a set of intensity signals corresponding to the input sample based on the reference signal distribution, determining, for the one or more input samples, and from a respective one or more calibrated sets of intensity signals corresponding to the one or more input samples, a respective at least one aggregated calibrated signal from targeted genomic region(s) to produce a collection of aggregated calibrated signals, and detecting variant(s) in the targeted genomic region(s) based on the collection of aggregated calibrated signals.
Need to check novelty before this filing date? Find Prior Art

Description

Abstract: Targeted copy number detection based on a wafer, such as detection of contaminated and / or unevenly concentrated samples, includes: acquiring a set of intensity signals from the detection of a set of input samples; performing cross-sample calibration on these intensity signals based on a reference sample, the calibration including: constructing a reference signal distribution based on the intensity signals of the reference sample; calibrating a set of intensity signals corresponding to one or more input samples based on the reference signal distribution; for one or more input samples, determining at least one aggregated calibration signal from a target genomic region from one or more calibrated intensity signal sets corresponding to the one or more input samples, thereby generating an aggregated calibration signal set; and detecting variations in the target genomic region based on the aggregated calibration signal set.

Claims

CLAIMSWhat is claimed is:

1. A computer-implemented method comprising: obtaining a collection of intensity signals from assays of a set of input samples comprising genetic material; performing a cross-sample calibration on the intensity signals of the collection of intensity signals based on one or more reference samples, the performing the cross-sample calibration comprising: constructing a reference signal distribution based on intensity signals of the one or more reference samples; and for one or more input samples of the set of input samples: obtaining a respective set of intensity signals, of the collection of intensity signals, corresponding to that input sample, the set of intensity signals corresponding to the input sample comprising (i) a first subset, C, of intensity signals from one or more targeted genomic regions of interest and (ii) a second subset, , of intensity signals from at least one genomic regions outside the one or more targeted genomic regions of interest; and calibrating the intensity signals in C based on the reference signal distribution, to produce a respective calibrated set of intensity signals corresponding to the input sample; determining, for the one or more input samples, and from a respective one or more calibrated sets of intensity signals corresponding to the one or more input samples, a respective at least one aggregated calibrated signal from the one or more targeted genomic regions of interest, wherein the determining produces a collection of aggregated calibrated signals; anddetecting one or more variants in the one or more targeted genomic regions of interest based on the collection of aggregated calibrated signals.

2. The method of claim 1, wherein the calibrating of the intensity signals in C, of the set of intensity signals corresponding to the input sample, comprises building a mapping for that input sample based on relations between (i) the intensity signals in and (ii) the reference signal distribution.

3. The method of claim 2, wherein the building the mapping comprises defining a mapping function M(x) such that M(x) maps intensity signal x as: for x existing in B, M(x) = a matching intensity signal from a vector, A, of reference signal intensities, from the reference signal distribution, corresponding to the at least one genomic regions outside the one or more targeted genomic regions of interest; for x not existing in B but falling between multiple intensity signals in , M(x) = a linear interpolation based on the M(x) mappings of the multiple intensity signals in ; and for x not existing in B and not falling within a range of the intensity signals in , M(x) = an extrapolation based on mappings of highest and lowest quantiles in B.

4. The method of claim 3, wherein the constructing the reference signal distribution computes the vector^ as cross-sample medians of autosomal array probes that are outside the one or more targeted genomic regions of interest.

5. The method of claims 3 or 4, wherein the calibrating the intensity signals in C further comprises using the mapping function to map the intensity signals in C to produce the calibrated set of intensity signals corresponding to the input sample.

6. The method of claim 1, 2, 3, 4, or 5, wherein the obtaining the collection of intensity signals comprises, for the set of input samples, using a set of array hybridization control probes to identify probe hybridization biases by aggregating row-based normalized raw intensity values from the control probes intoan aggregated value cs, aggregating row-based normalized intensity values from assays targeting human genomic material into an aggregated value xs, and determining a contamination factor fsas a function of xsand cs, where fs, xsand csare determined per input sample.

7. The method of claim 6, wherein the function for contamination factor fsis. fs Xs / s-8. The method of claim 6 or 7, wherein the determining, for the one or more input samples, and from the respective one or more calibrated sets of intensity signals corresponding to the one or more input samples, the respective at least one aggregated calibrated signal comprises, for an aggregated calibrated signal of the at least one aggregated calibrated signal: determining a first aggregated signal from a calibrated set of intensity signals corresponding to a targeted region of the input sample; and using the contamination factor to correct the first aggregated signal and produce a second aggregated signal, wherein the second aggregated signal is output as the aggregated calibrated signal for the targeted region of the input sample.

9. The method of claim 8, wherein the using the contamination factor and producing the second aggregated signal comprises (i) using a regression-based model to predict contribution of contamination based on the contamination factor, (ii) determining a residue as a function of the first aggregated signal and the contribution of contamination predicted by the model, and (iii) determining the second aggregated signal as a function of the residue and a composite contamination factor from across the input samples.

10. The method of claim 1, 2, 3, 4, 5, 6, 7, 8, or 9, wherein the one or more variants are one or more copy number variants.

11. The method of claim 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, wherein none of (i) deoxyribonucleic acid (DNA) quantification of the input samples, (ii) normalization of the input samples, and (iii) prior measurements of fraction or amount of DNA contaminant in the input samples is known or required in performing the method.

12. The method of claim 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11, wherein the input samples of the set of input samples contains at least one of (i) variable amounts or concentrations of deoxyribonucleic acid (DNA) relative to each other or (ii) different fractions of contaminant DNA relative to each other.

13. The method of claim 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12, wherein the collection of intensity signals is from a high-throughput genotyping platform genotyping the input samples using a microarray -based genotyping platform.

14. A computer system comprising: a memory; and a processor in communication with the memory, wherein the computer system is configured to perform a method comprising: obtaining a collection of intensity signals from assays of a set of input samples comprising genetic material; performing a cross-sample calibration on the intensity signals of the collection of intensity signals based on one or more reference samples, the performing the cross-sample calibration comprising: constructing a reference signal distribution based on intensity signals of the one or more reference samples; and for one or more input samples of the set of input samples: obtaining a respective set of intensity signals, of the collection of intensity signals, corresponding to that input sample, the set of intensity signals corresponding to the input sample comprising (i) a first subset, C, of intensity signals from one or more targeted genomic regions of interest and (ii) a second subset, 7>, of intensity signals from at least one genomic regions outside the one or more targeted genomic regions of interest; andcalibrating the intensity signals in C based on the reference signal distribution, to produce a respective calibrated set of intensity signals corresponding to the input sample; determining, for the one or more input samples, and from a respective one or more calibrated sets of intensity signals corresponding to the one or more input samples, a respective at least one aggregated calibrated signal from the one or more targeted genomic regions of interest, wherein the determining produces a collection of aggregated calibrated signals; and detecting one or more variants in the one or more targeted genomic regions of interest based on the collection of aggregated calibrated signals.

15. The computer system of claim 14, wherein the calibrating of the intensity signals in C, of the set of intensity signals corresponding to the input sample, comprises building a mapping for that input sample based on relations between (i) the intensity signals in B and (ii) the reference signal distribution.

16. The computer system of claim 15, wherein the building the mapping comprises defining a mapping function M(x) such that M(x) maps intensity signal x as: for x existing in B, M(x) = a matching intensity signal from a vector, A, of reference signal intensities, from the reference signal distribution, corresponding to the at least one genomic regions outside the one or more targeted genomic regions of interest; for x not existing in B but falling between multiple intensity signals in , M(x) = a linear interpolation based on the M(x) mappings of the multiple intensity signals in ; and for x not existing in B and not falling within a range of the intensity signals in , M(x) = an extrapolation based on mappings of highest and lowest quantiles in B.

17. The computer system of claim 16, wherein the constructing the reference signal distribution computes the vector^ as cross-sample medians of autosomal array probes that are outside the one or more targeted genomic regions of interest.

18. The computer system of claim 16 or 17, wherein the calibrating the intensity signals in C further comprises using the mapping function to map the intensity signals in C to produce the calibrated set of intensity signals corresponding to the input sample.

19. The computer system of claim 14, 15, 16, 17, or 18, wherein the obtaining the collection of intensity signals comprises, for the set of input samples, using a set of array hybridization control probes to identify probe hybridization biases by aggregating row-based normalized raw intensity values from the control probes into an aggregated value cs, aggregating row-based normalized intensity values from assays targeting human genomic material into an aggregated value xs, and determining a contamination factor fsas a function of xsand cs, where fs, xsand csare determined per input sample.

20. The computer system of claim 19, wherein the function for contamination factor fsis: fs= xs / cs.

21. The computer system of claim 19 or 20, wherein the determining, for the one or more input samples, and from the respective one or more calibrated sets of intensity signals corresponding to the one or more input samples, the respective at least one aggregated calibrated signal comprises, for an aggregated calibrated signal of the at least one aggregated calibrated signal: determining a first aggregated signal from a calibrated set of intensity signals corresponding to a targeted region of the input sample; and using the contamination factor to correct the first aggregated signal and produce a second aggregated signal, wherein the second aggregated signal is output as the aggregated calibrated signal for the targeted region of the input sample.

22. The computer system of claim 21, wherein the using the contamination factor and producing the second aggregated signal comprises (i) using a regressionbased model to predict contribution of contamination based on the contamination factor, (ii) determining a residue as a function of the first aggregated signal and the contribution of contamination predicted by the model, and (iii) determining the second aggregated signal as a function of the residue and a composite contamination factor from across the input samples.

23. The computer system of claim 14, 15, 16, 17, 18, 19, 20, 21, or 22, wherein the one or more variants are one or more copy number variants.

24. The computer system of claim 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23, wherein none of (i) deoxyribonucleic acid (DNA) quantification of the input samples, (ii) normalization of the input samples, and (iii) prior measurements of fraction or amount of DNA contaminant in the input samples is known or required in performing the method.

25. The computer system of claim 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24, wherein the input samples of the set of input samples contains at least one of (i) variable amounts or concentrations of deoxyribonucleic acid (DNA) relative to each other or (ii) different fractions of contaminant DNA relative to each other.

26. The computer system of claim 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25, wherein the collection of intensity signals is from a high-throughput genotyping platform genotyping the input samples using a microarray -based genotyping platform.

27. A computer program product comprising: a computer readable storage medium readable by a processing circuit and storing instructions for execution by the processing circuit for performing a method comprising: obtaining a collection of intensity signals from assays of a set of input samples comprising genetic material;performing a cross-sample calibration on the intensity signals of the collection of intensity signals based on one or more reference samples, the performing the cross-sample calibration comprising: constructing a reference signal distribution based on intensity signals of the one or more reference samples; and for one or more input samples of the set of input samples: obtaining a respective set of intensity signals, of the collection of intensity signals, corresponding to that input sample, the set of intensity signals corresponding to the input sample comprising (i) a first subset, C, of intensity signals from one or more targeted genomic regions of interest and (ii) a second subset, , of intensity signals from at least one genomic regions outside the one or more targeted genomic regions of interest; and calibrating the intensity signals in C based on the reference signal distribution, to produce a respective calibrated set of intensity signals corresponding to the input sample; determining, for the one or more input samples, and from a respective one or more calibrated sets of intensity signals corresponding to the one or more input samples, a respective at least one aggregated calibrated signal from the one or more targeted genomic regions of interest, wherein the determining produces a collection of aggregated calibrated signals; and detecting one or more variants in the one or more targeted genomic regions of interest based on the collection of aggregated calibrated signals.

28. The computer program product of claim 27, wherein the calibrating of the intensity signals in C, of the set of intensity signals corresponding to the input sample, comprises building a mapping for that input sample based on relations between (i) the intensity signals in B and (ii) the reference signal distribution.

29. The computer program product of claim 28, wherein the building the mapping comprises defining a mapping function M(x) such that M(x) maps intensity signal x as: for x existing in B, M(x) = a matching intensity signal from a vector, A, of reference signal intensities, from the reference signal distribution, corresponding to the at least one genomic regions outside the one or more targeted genomic regions of interest; for x not existing in B but falling between multiple intensity signals in , M(x) = a linear interpolation based on the M(x) mappings of the multiple intensity signals in ; and for x not existing in B and not falling within a range of the intensity signals in , M(x) = an extrapolation based on mappings of highest and lowest quantiles in B.

30. The computer program product of claim 29, wherein the constructing the reference signal distribution computes the vector^ as cross-sample medians of autosomal array probes that are outside the one or more targeted genomic regions of interest.

31. The computer program product of claim 29 or 30, wherein the calibrating the intensity signals in C further comprises using the mapping function to map the intensity signals in C to produce the calibrated set of intensity signals corresponding to the input sample.

32. The computer program product of claim 27, 28, 29, 30, or 31, wherein the obtaining the collection of intensity signals comprises, for the set of input samples, using a set of array hybridization control probes to identify probe hybridization biases by aggregating row-based normalized raw intensity values from the control probes into an aggregated value cs, aggregating row-based normalizedintensity values from assays targeting human genomic material into an aggregated value xs, and determining a contamination factor fsas a function of xsand cs, where fs, xsand csare determined per input sample.

33. The computer program product of claim 32, wherein the function for contamination factor fsis: fs= xs / cs.

34. The computer program product of claim 32 or 33, wherein the determining, for the one or more input samples, and from the respective one or more calibrated sets of intensity signals corresponding to the one or more input samples, the respective at least one aggregated calibrated signal comprises, for an aggregated calibrated signal of the at least one aggregated calibrated signal: determining a first aggregated signal from a calibrated set of intensity signals corresponding to a targeted region of the input sample; and using the contamination factor to correct the first aggregated signal and produce a second aggregated signal, wherein the second aggregated signal is output as the aggregated calibrated signal for the targeted region of the input sample.

35. The computer program product of claim 34, wherein the using the contamination factor and producing the second aggregated signal comprises (i) using a regression-based model to predict contribution of contamination based on the contamination factor, (ii) determining a residue as a function of the first aggregated signal and the contribution of contamination predicted by the model, and (iii) determining the second aggregated signal as a function of the residue and a composite contamination factor from across the input samples.

36. The computer program product of claim 27, 28, 29, 30, 31, 32, 33, 34, or 35, wherein the one or more variants are one or more copy number variants.

37. The computer program product of claim 27, 28, 29, 30, 31, 32, 33, 34, 35, or 36, wherein none of (i) deoxyribonucleic acid (DNA) quantification of the input samples, (ii) normalization of the input samples, and (iii) prior measurements of fraction or amount of DNA contaminant in the input samples is known or required in performing the method.

38. The computer program product of claim 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, or 37, wherein the input samples of the set of input samples contains at least one of (i) variable amounts or concentrations of deoxyribonucleic acid (DNA) relative to each other or (ii) different fractions of contaminant DNA relative to each other.

39. The computer program product of claim 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, or 38, wherein the collection of intensity signals is from a high-throughput genotyping platform genotyping the input samples using a microarray -based genotyping platform.

40. A computer-implemented method comprising: obtaining a collection of intensity signals from assays of a set of input samples comprising genetic material; using a set of array hybridization control probes to identify probe hybridization biases by aggregating row-based normalized raw intensity values from the control probes into an aggregated value cs, aggregating rowbased normalized intensity values from assays targeting human genomic material into an aggregated value xs, and determining a contamination factor as a function of xsand cs, where fs, xsand csare determined per input sample of the set of input samples; and using the contamination factor to correct a first signal obtained based on intensity signals of the collection of intensity signals and produce a corrected signal.

41. The method of claim 40, wherein using the contamination factor and producing the corrected signal comprises (i) using a regression-based model to predict contribution of contamination based on the contamination factor, (ii) determining a residue as a function of the first signal and the contribution of contamination predicted by the model, and (iii) determining the corrected signal as a function of the residue and a composite contamination factor from across the input samples.

42. The method of claim 40 or 41, wherein the function for contamination factor fsis: fs= xs / cs.

43. The method of claim 40, 41, or 42, wherein the first signal is a first aggregated signal from a set of the intensity signals of the collection of intensity signals, the first aggregated signal corresponding to a target region of an input sample of the set of input samples, and wherein the corrected signal is a corrected aggregated signal.

44. The method of claim 40, 41, 42, or 43, wherein the collection of intensity signals is from a high-throughput genotyping platform genotyping the input samples using a microarray -based genotyping platform.

45. A computer system comprising: a memory; and a processor in communication with the memory, wherein the computer system is configured to perform a method comprising: obtaining a collection of intensity signals from assays of a set of input samples comprising genetic material; using a set of array hybridization control probes to identify probe hybridization biases by aggregating row-based normalized raw intensity values from the control probes into an aggregated value cs, aggregating row-based normalized intensity values from assays targeting human genomic material into an aggregated value xs, and determining a contamination factor fsas a function of xsand cs, where fs, xsand csare determined per input sample of the set of input samples; and using the contamination factor to correct a first signal obtained based on intensity signals of the collection of intensity signals and produce a corrected signal.

46. The computer system of claim 45, wherein using the contamination factor and producing the corrected signal comprises (i) using a regression-based model to predict contribution of contamination based on the contamination factor, (ii) determining a residue as a function of the first signal and the contribution ofcontamination predicted by the model, and (iii) determining the corrected signal as a function of the residue and a composite contamination factor from across the input samples.

47. The computer system of claim 45 or 46, wherein the function for contamination factor fsis: fs= xs / cs.

48. The computer system of claim 45, 46, or 47, wherein the first signal is a first aggregated signal from a set of the intensity signals of the collection of intensity signals, the first aggregated signal corresponding to a target region of an input sample of the set of input samples, and wherein the corrected signal is a corrected aggregated signal.

49. The computer system of claim 45, 46, 47, or 48, wherein the collection of intensity signals is from a high-throughput genotyping platform genotyping the input samples using a microarray-based genotyping platform.

50. A computer program product comprising: a computer readable storage medium readable by a processing circuit and storing instructions for execution by the processing circuit for performing a method comprising: obtaining a collection of intensity signals from assays of a set of input samples comprising genetic material; using a set of array hybridization control probes to identify probe hybridization biases by aggregating row-based normalized raw intensity values from the control probes into an aggregated value cs, aggregating row-based normalized intensity values from assays targeting human genomic material into an aggregated value xs, and determining a contamination factor as a function of xsand cs, where fs, xsand csare determined per input sample of the set of input samples; andusing the contamination factor to correct a first signal obtained based on intensity signals of the collection of intensity signals and produce a corrected signal.

51. The computer program product of claim 50, wherein using the contamination factor and producing the corrected signal comprises (i) using a regression-based model to predict contribution of contamination based on the contamination factor, (ii) determining a residue as a function of the first signal and the contribution of contamination predicted by the model, and (iii) determining the corrected signal as a function of the residue and a composite contamination factor from across the input samples.

52. The computer program product of claim 50 or 51, wherein the function for contamination factor / s: fs= xs / cs.

53. The computer program product of claim 50, 51, or 52, wherein the first signal is a first aggregated signal from a set of the intensity signals of the collection of intensity signals, the first aggregated signal corresponding to a target region of an input sample of the set of input samples, and wherein the corrected signal is a corrected aggregated signal.

54. The computer program product of claim 50, 51, 52, or 53, wherein the collection of intensity signals is from a high-throughput genotyping platform genotyping the input samples using a microarray -based genotyping platform.